Bookmarks

Bookmarks

49363 bookmarks
Custom sorting
freee請求書 33,000 行の openapi.yml を TypeSpec に移行した話 - freee Developers Hub
freee請求書 33,000 行の openapi.yml を TypeSpec に移行した話 - freee Developers Hub
債権・請求書ドメインでエンジニアをやっている jaxx です。 好きな ESP32 デバイスは M5Stack ATOMS3 Lite です。 freee請求書では、Public API だけでなく、内部的なAPIも含めると複数あり、最も多いAPI群は、生成後のOpenAPIが 33,000行ありましたが、今回 API…
·developers.freee.co.jp·
freee請求書 33,000 行の openapi.yml を TypeSpec に移行した話 - freee Developers Hub
(1)ステラ・ビダーマンのXに関する発言:「現時点では、@SakanaAILabsの主張は一切信用すべきではありません。このツイートにある明白な虚偽の主張と、昨年発表した一連の不正な『AI科学者』論文を考えると、彼らの言うことの正反対が真実であるというのが私の基本的な推測です。」/Twitter
(1)ステラ・ビダーマンのXに関する発言:「現時点では、@SakanaAILabsの主張は一切信用すべきではありません。このツイートにある明白な虚偽の主張と、昨年発表した一連の不正な『AI科学者』論文を考えると、彼らの言うことの正反対が真実であるというのが私の基本的な推測です。」/Twitter
·x.com·
(1)ステラ・ビダーマンのXに関する発言:「現時点では、@SakanaAILabsの主張は一切信用すべきではありません。このツイートにある明白な虚偽の主張と、昨年発表した一連の不正な『AI科学者』論文を考えると、彼らの言うことの正反対が真実であるというのが私の基本的な推測です。」/Twitter
AIのモデル崩壊と多様性 - ジョイジョイジョイ
AIのモデル崩壊と多様性 - ジョイジョイジョイ
モデル崩壊 (model collapse) とは、AI が出力したテキストや画像がインターネットにアップロードされ、それらが AI の訓練データに混入し、そのデータで訓練された AI が出力したものがアップロードされる、というサイクルを繰り返すことで AI の性能が崩壊することです。最も有名な研究は Nature に…
·joisino.hatenablog.com·
AIのモデル崩壊と多様性 - ジョイジョイジョイ
エージェントスキルを評価する仕組みを作ってみる | フューチャー技術ブログ
エージェントスキルを評価する仕組みを作ってみる | フューチャー技術ブログ
コーディングエージェントの様々なツール横断で使えるハーネスといえば、skillsという雰囲気になってきました。AGENTS.mdやMCPもありますが、作りやすさや小回りが効く点、プロジェクトやチーム単位で気軽に改善していける点など、人気になるのはうなづけます。もちろん、MCPが全てにおいて劣っているというわけではないので、必要に応じて使い分けることになると思います。 項目 AGENTS.md
·future-architect.github.io·
エージェントスキルを評価する仕組みを作ってみる | フューチャー技術ブログ
Cloudflare の一時アカウントを使って即座にデプロイできるようになった
Cloudflare の一時アカウントを使って即座にデプロイできるようになった
Cloudflare の Temporary Cloudflare Accounts を使用すると、人間が介入することなく AI エージェントが即座に Cloudflare Workers にデプロイできるようになります。この記事では、Temporary Cloudflare Accounts を使用して実際に Cloudflare Workers にデプロイする方法を試してみます。
·azukiazusa.dev·
Cloudflare の一時アカウントを使って即座にデプロイできるようになった
ワークフローを再利用可能なスキルに変換する Record & Replay を試してみた
ワークフローを再利用可能なスキルに変換する Record & Replay を試してみた
Codex の Record & Replay は macOS 上でのユーザーの操作を実演することで再利用可能なスキルに変換する機能です。例えば経費精算の提出や勤怠アプリへの打刻や工数入力、定期的なレポートの作成などをスキルとして記録し、煩雑な定型業務を AI に任せることが期待できます。この記事では、Record & Replay を実際に試してみた様子を紹介します。
·azukiazusa.dev·
ワークフローを再利用可能なスキルに変換する Record & Replay を試してみた
Ponytail? YAGNI!
Ponytail? YAGNI!
The post examines Ponytail, a popular AI coding “skill”, and argues that its benchmarked benefits appear to come largely from encouraging terse, YAGNI-style responses rather than from any deeper engineering value. By showing that a simple prompt can match or beat Ponytail on its own benchmark, it makes a broader case for treating prompt-based tools with scepticism unless their claims are backed by robust evaluation.
·blog.scottlogic.com·
Ponytail? YAGNI!
AI エージェントフレームワーク eve を試してみた
AI エージェントフレームワーク eve を試してみた
Vercel が新しい AI エージェントフレームワーク eve を発表しました。Next.js の設計思想に基づいて構築された eve は、AI エージェントの開発に必要な機能がすべて揃ったフレームワークです。この記事では、eve を使って簡単なエージェントを作成し、実行する方法を紹介します。
·azukiazusa.dev·
AI エージェントフレームワーク eve を試してみた
eve – The Agent Framework - Vercel
eve – The Agent Framework - Vercel
Like Next.js for web apps, but for agents. Markdown for instructions and skills, TypeScript for tools. Durable by default.
·vercel.com·
eve – The Agent Framework - Vercel
Artificial Analysis on X: "Announcing AA-Briefcase, the benchmark for the next era of agentic knowledge work AA-Briefcase is our new benchmark for testing models on long-horizon knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week projects, each with many https://t.co/s1LJxN2Nct" / Twitter
Artificial Analysis on X: "Announcing AA-Briefcase, the benchmark for the next era of agentic knowledge work AA-Briefcase is our new benchmark for testing models on long-horizon knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week projects, each with many https://t.co/s1LJxN2Nct" / Twitter
AA-Briefcase is our new benchmark for testing models on long-horizon knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week projects, each with many
·x.com·
Artificial Analysis on X: "Announcing AA-Briefcase, the benchmark for the next era of agentic knowledge work AA-Briefcase is our new benchmark for testing models on long-horizon knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week projects, each with many https://t.co/s1LJxN2Nct" / Twitter
社内にHTMLをホストする環境を作ったら社内情報の流れが変わった - BASEプロダクトチームブログ
社内にHTMLをホストする環境を作ったら社内情報の流れが変わった - BASEプロダクトチームブログ
CTO の川口(id:dmnlk)です。 最近、社内で「とりあえず HTML にして置いておくね」という会話を当たり前のように耳にするようになりました。きっかけは、社内 HTML をホストするだけのささやかな環境をひとつ用意したことです。たったそれだけのことなのに、気づけばエンジニア以外のメンバーまで使い始め、社内の情…
·devblog.thebase.in·
社内にHTMLをホストする環境を作ったら社内情報の流れが変わった - BASEプロダクトチームブログ
InsForge/InsForge: The all-in-one, open-source backend platform for agentic coding. InsForge gives your coding agent database, auth, storage, compute, hosting, and AI gateway to ship full-stack apps end-to-end.
InsForge/InsForge: The all-in-one, open-source backend platform for agentic coding. InsForge gives your coding agent database, auth, storage, compute, hosting, and AI gateway to ship full-stack apps end-to-end.
The all-in-one, open-source backend platform for agentic coding. InsForge gives your coding agent database, auth, storage, compute, hosting, and AI gateway to ship full-stack apps end-to-end. - Ins...
·github.com·
InsForge/InsForge: The all-in-one, open-source backend platform for agentic coding. InsForge gives your coding agent database, auth, storage, compute, hosting, and AI gateway to ship full-stack apps end-to-end.
InsForge - The agent-native cloud infrastructure platform
InsForge - The agent-native cloud infrastructure platform
Model gateway, compute, deployment, database, auth, and more. Every service built for AI coding agents to operate end to end through CLI and skills.
·insforge.dev·
InsForge - The agent-native cloud infrastructure platform
The Reflect Package | Internals for Interns
The Reflect Package | Internals for Interns
In the previous article we watched the runtime rebuild an entire stack trace out of metadata the compiler and linker had frozen into the binary at build time. I told you at the end that reflect works on exactly the same trick — metadata baked into the binary, only pointed at your data instead of your call stack. Today we’re going to cash that promise in. Let’s start with a program that, the first time you see it, feels like it shouldn’t be possible:
·internals-for-interns.com·
The Reflect Package | Internals for Interns