Bookmarks

Bookmarks

49953 bookmarks
Custom sorting
Where Does current_user Actually Live?
Where Does current_user Actually Live?
The request-local life of identity across Warden, Devise, CurrentAttributes, and the Rails Executor.
·railsrevelry.substack.com·
Where Does current_user Actually Live?
ts-bench v2: 数十万行規模のTypeScript製アプリの修正タスクでコーディングエージェントの性能を測る
ts-bench v2: 数十万行規模のTypeScript製アプリの修正タスクでコーディングエージェントの性能を測る
筆者は以前よりts-benchというAIコーディングエージェントのベンチマークを作っています。 このたびv2に移行したので、その背景と設計、そして今後どう継続していくかを書きます。 GitHub - laiso/ts-bench: Benchmark CLI for comparing AI coding agents on TypeScript workloads.Benchmark CLI for comparing AI coding agents on TypeScript workloads. - laiso/ts-benchGitHublaiso なお、本記事に出てくるスコアやティアは2026年4月に試験的に回した一回分のスナップショットです。 最新ではないし、モデルやエージェント製品、評価コードが更新されれば順位は動きます。 「Opusには解けず、Sonnetには解けた課題があった」といった観測も、現時点では結論ではなく途中経過として読んでください ハーネスという用語の整理 この記事では、エージェントハーネスを、大規模言語モデル(LLM)にプロンプトを渡し、ツー
·blog.lai.so·
ts-bench v2: 数十万行規模のTypeScript製アプリの修正タスクでコーディングエージェントの性能を測る
HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval...
HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval...
With the rapid spread of retrieval-augmented generation and semantic search, choosing the right embedding and retrieval configuration is increasingly hard. Large retrieval benchmarks are...
·arxiv.org·
HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval...
tpac_study (2026)
tpac_study (2026)
#tpac_study - https://web-study.connpass.com/event/378948/ X - https://twitter.com/sajikix
·speakerdeck.com·
tpac_study (2026)
WebMCP: Optimize Your Website for AI Agents | DebugBear
WebMCP: Optimize Your Website for AI Agents | DebugBear
Learn how WebMCP enables AI agents to interact with websites through structured tools instead of traditional browser automation. This guide explains how WebMCP works, how it compares to DOM-based approaches, and how to validate your WebMCP implementation using Lighthouse and DebugBear.
·debugbear.com·
WebMCP: Optimize Your Website for AI Agents | DebugBear
プロンプトは人手チューニングからAIチューニングへ:遺伝的アルゴリズムで回す自動最適化と高速化
プロンプトは人手チューニングからAIチューニングへ:遺伝的アルゴリズムで回す自動最適化と高速化
LINEヤフーの技術カンファレンス「Tech-Verse 2026」の公式記事です。こんにちは。LINEヤフー株式会社の中野です。Yahoo!検索のAI回答サービスで大規模言語モデル(LLM)の最適化...
·techblog.lycorp.co.jp·
プロンプトは人手チューニングからAIチューニングへ:遺伝的アルゴリズムで回す自動最適化と高速化
Hotswap - Drop-in open coding models hosted for you
Hotswap - Drop-in open coding models hosted for you
A drop-in API for open-source coding models. Keep Claude Code, Codex, and the tooling you already use. Pay as you go, cloud or dedicated infra, regional-hosting, optimized for performance, one-click model swap.
·hotswap.arcjet.com·
Hotswap - Drop-in open coding models hosted for you
TabularisDB/tabularis: Open-source database client for PostgreSQL, MySQL/MariaDB and SQLite with SQL notebooks, visual EXPLAIN, AI and MCP built in. Hackable with plugins.
TabularisDB/tabularis: Open-source database client for PostgreSQL, MySQL/MariaDB and SQLite with SQL notebooks, visual EXPLAIN, AI and MCP built in. Hackable with plugins.
Open-source database client for PostgreSQL, MySQL/MariaDB and SQLite with SQL notebooks, visual EXPLAIN, AI and MCP built in. Hackable with plugins. - TabularisDB/tabularis
·github.com·
TabularisDB/tabularis: Open-source database client for PostgreSQL, MySQL/MariaDB and SQLite with SQL notebooks, visual EXPLAIN, AI and MCP built in. Hackable with plugins.
GitHub - omnigent-ai/omnigent: Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.
GitHub - omnigent-ai/omnigent: Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.
Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, an...
·github.com·
GitHub - omnigent-ai/omnigent: Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.
Omnigent — a meta-harness for building and running AI agents
Omnigent — a meta-harness for building and running AI agents
Omnigent is an open-source meta-harness: a common layer for composing, governing, and collaborating on AI agents — coding and otherwise — on top of the harnesses you already use.
·omnigent.ai·
Omnigent — a meta-harness for building and running AI agents
tk0miya/rubocop-rbs_inline
tk0miya/rubocop-rbs_inline
Contribute to tk0miya/rubocop-rbs_inline development by creating an account on GitHub.
·github.com·
tk0miya/rubocop-rbs_inline
Astro 7.0 | Astro
Astro 7.0 | Astro
Astro 7.0 brings faster builds with Vite 8, a new Rust compiler, Advanced Routing, background dev server support, and structured logging.
·astro.build·
Astro 7.0 | Astro
freee請求書 33,000 行の openapi.yml を TypeSpec に移行した話 - freee Developers Hub
freee請求書 33,000 行の openapi.yml を TypeSpec に移行した話 - freee Developers Hub
債権・請求書ドメインでエンジニアをやっている jaxx です。 好きな ESP32 デバイスは M5Stack ATOMS3 Lite です。 freee請求書では、Public API だけでなく、内部的なAPIも含めると複数あり、最も多いAPI群は、生成後のOpenAPIが 33,000行ありましたが、今回 API…
·developers.freee.co.jp·
freee請求書 33,000 行の openapi.yml を TypeSpec に移行した話 - freee Developers Hub
(1)ステラ・ビダーマンのXに関する発言:「現時点では、@SakanaAILabsの主張は一切信用すべきではありません。このツイートにある明白な虚偽の主張と、昨年発表した一連の不正な『AI科学者』論文を考えると、彼らの言うことの正反対が真実であるというのが私の基本的な推測です。」/Twitter
(1)ステラ・ビダーマンのXに関する発言:「現時点では、@SakanaAILabsの主張は一切信用すべきではありません。このツイートにある明白な虚偽の主張と、昨年発表した一連の不正な『AI科学者』論文を考えると、彼らの言うことの正反対が真実であるというのが私の基本的な推測です。」/Twitter
·x.com·
(1)ステラ・ビダーマンのXに関する発言:「現時点では、@SakanaAILabsの主張は一切信用すべきではありません。このツイートにある明白な虚偽の主張と、昨年発表した一連の不正な『AI科学者』論文を考えると、彼らの言うことの正反対が真実であるというのが私の基本的な推測です。」/Twitter
AIのモデル崩壊と多様性 - ジョイジョイジョイ
AIのモデル崩壊と多様性 - ジョイジョイジョイ
モデル崩壊 (model collapse) とは、AI が出力したテキストや画像がインターネットにアップロードされ、それらが AI の訓練データに混入し、そのデータで訓練された AI が出力したものがアップロードされる、というサイクルを繰り返すことで AI の性能が崩壊することです。最も有名な研究は Nature に…
·joisino.hatenablog.com·
AIのモデル崩壊と多様性 - ジョイジョイジョイ
エージェントスキルを評価する仕組みを作ってみる | フューチャー技術ブログ
エージェントスキルを評価する仕組みを作ってみる | フューチャー技術ブログ
コーディングエージェントの様々なツール横断で使えるハーネスといえば、skillsという雰囲気になってきました。AGENTS.mdやMCPもありますが、作りやすさや小回りが効く点、プロジェクトやチーム単位で気軽に改善していける点など、人気になるのはうなづけます。もちろん、MCPが全てにおいて劣っているというわけではないので、必要に応じて使い分けることになると思います。 項目 AGENTS.md
·future-architect.github.io·
エージェントスキルを評価する仕組みを作ってみる | フューチャー技術ブログ
Cloudflare の一時アカウントを使って即座にデプロイできるようになった
Cloudflare の一時アカウントを使って即座にデプロイできるようになった
Cloudflare の Temporary Cloudflare Accounts を使用すると、人間が介入することなく AI エージェントが即座に Cloudflare Workers にデプロイできるようになります。この記事では、Temporary Cloudflare Accounts を使用して実際に Cloudflare Workers にデプロイする方法を試してみます。
·azukiazusa.dev·
Cloudflare の一時アカウントを使って即座にデプロイできるようになった