GPTProto

Tutorials

434 picksNewest firstLatest article Oct 2, 2026, 8:00 AM
2 stories
6 stories
20:00 Anthropic:Claude.dev 开发者博客(RSS)AI score 68/100

Getting started with Claude Code mods

Claude Code already lets you change a lot about how it behaves: settings, permission rules, slash commands, skills and a status line. Mods go further. Mods can rewrite or replace what Claude Code does, and can even draw custom UI. Under the hood, mods are hooks, and they ship inside plugins. Each one is a small JavaScript or TypeScript module that runs inside your session and sees every event as it happens. That makes mods a way to fit Claude Code to how you work. You can add a readout you check

18:54 Ethan Mollick:One Useful Thing(RSS)AI score 64/100

The Dot and the Swarm

I generally think I have done a good job anticipating the direction and pace of AI over the few years I have been writing this Substack, but I think I recently got something fairly large wrong. In the last year I have been posting about how I suspected that humans would have to approach working with agents as a manager, deciding how to delegate work to agents and specifying how those agents should be organized. I thought that getting agents to work effectively as a group would take careful const

08:00 OpenRouter:Announcements(RSS)AI score 68/100

Confidence Thresholds for Model Escalation Routing

Sending every request to your strongest model gets good answers at the highest price, because you pay frontier rates for requests a cheaper model would have answered correctly. Fixed routing rules, such as picking the model by keyword or task type, cost less but need rewriting as your traffic changes. Confidence-based escalation sits between the two. You ask the model to score its own answer, then route on that score. Answers with a high score stay on a cheap model. Answers with a low score go t

08:00 OpenRouter:Announcements(RSS)AI score 67/100

How to Gate Pull Requests on LLM Evals in CI

Changing one line in a support agent’s system prompt can ship an agent that tells customers the refund window is 30 days when your policy says 14. Nothing in a normal CI pipeline checks what the model says, so the build passes and the first person to see the wrong answer is a customer.Gating a pull request on a fixed eval set works the same way as gating on a failing unit test. You keep test cases in the repository, run them when a prompt changes, and block the merge when too many fail.In this g

08:00 OpenRouter:Announcements(RSS)AI score 67/100

Cost vs. Quality Tradeoff Framework for Agent Models

Picking a model for an agent by leaderboard rank pays frontier prices for tasks that a cheaper model may handle at the same accuracy. The question to answer is not which model scores highest, but which model is the cheapest one that is good enough for the task in front of you. This guide is a three-step framework for making that call. You set the quality bar the task needs, measure cost per quality point on your own examples, and pick the cheapest model that clears the bar with margin. Tl;dr Set

07:00 英国 AI Security Institute:Blog(网页)AI score 70/100

Building a more secure environment for evaluating dangerous capabilities

In August, we reported an incident in which AI agents, during a cyber evaluation, took sustained action against real people beyond the remit of their task. Our incident was one of several across the sector in which AI agents took actions during evaluations that their operators had not intended. Although the circumstances differed, these incidents highlight the need to ensure that there are robust security practices underpinning frontier AI development and research. In response, we paused our hig

4 stories
08:00 OpenRouter:Announcements(RSS)AI score 65/100

Building a Golden Eval Dataset from Production Traffic

You update a prompt, or a provider rolls out a new checkpoint under the same model ID, and something in production regresses. The last week of user complaints looks slightly different from the week before. A public benchmark like MMLU won’t catch that. It measures general capability across academic subjects, not how the model handles your product’s traffic. A golden eval dataset closes that gap. It’s a curated collection of production inputs paired with reviewed expected outputs, versioned in Gi

08:00 OpenRouter:Announcements(RSS)AI score 69/100

AI Agent Regression Testing After a Prompt or Model Change

A ~author/family-latest alias always resolves to the newest concrete model in a family. That’s convenient in production and a problem in a regression test, because the model can change between runs without any change in your repository. Our latest model resolution docs describe the mechanism and recommend a concrete model slug when you need a fixed version for reproducibility. This guide covers the locked case set and the behavioral contract, then the model-swap case in detail. Tl;dr Regression

08:00 OpenRouter:Announcements(RSS)AI score 62/100

How to Test Tool-Calling Accuracy in AI Agents

An agent can fail in two places when it uses a tool. It can choose the wrong tool, or it can choose the right one and send the wrong arguments. Those failures tell you different things. If an agent calls lookup_order instead of refund_order, the problem is tool selection. If it calls refund_order with the wrong order_id, it chose the right tool and passed the wrong arguments. This guide covers three ways to test tool-calling behavior. The first is a reference-free large language model (LLM) judg

00:00 METR:Blog(网页)AI score 72/100

Chris Painter's testimony to the U.S. Senate on AI agent incidents

On September 30, 2026, METR President Chris Painter testified before the U.S. Senate Committee on Homeland Security & Governmental Affairs’ Subcommittee on Disaster Management, District of Columbia, and Census, at a hearing titled “Rogue AI: Securing the Homeland Against AI Agent Attacks”. The written testimony is published in full below, and is also available as a PDF. Introduction Chairman Hawley, Ranking Member Kim, and members of the subcommittee, thank you for inviting me to testify today.

4 stories
17:00 Sarvam AI(网页)AI score 74/100

Building AI Agents: A First-Principles Guide

An agent is a language model running in a loop, with tools it can call and instructions that tell it how to behave. This guide builds up from that idea to a working support agent.ResearchSeptember 29, 2026·25 min readAn agent is a language model running in a loop, with tools it can call and instructions that tell it how to behave. Everything else - skills, memory, domain knowledge - is a way of putting the right text in front of the model at the right time.How to read thisThis guide goes top to

08:00 vLLM 官方博客(RSS)AI score 62/100

Taking vLLM Apart: A Practical Guide to Disaggregated Serving

A single vllm serve process does three jobs that get in each other's way: processing prompts (prefill), generating tokens (decode) and a pile of CPU work around them. Disaggregated serving in vLLM separates the different stages of LLM inference. Splitting prefill from decode stops long prompts stalling everyone else's output, as long as the KV cache moves between them fast. Moving tokenization and parsing to a CPU-only frontend (/render, /derender) takes that work off your GPU nodes and leaves t

04:11 Databricks:Blog(RSS)AI score 65/100

How Databricks rolls out frontier models to 12,000 employees on Day 1

Providing our employees access to frontier AI capabilities is a top priority at Databricks, and consequently, it is important to us for them to use new models instantly when they become available. At the same time, it is nontrivial to give more than 12,000 people rapid access to a new model because:Models that are marketed as frontier often aren’t. For example, Opus 5.0 was more expensive and ranked lower on both quantitative and qualitative quality scores among our engineers compared with Opus

03:00 GitHub BlogAI score 71/100

How we found 24 Android vulnerabilities using our open source AI security agent

With the rise of AI in the security space, our team created the GitHub Security Lab Taskflow Agent as a way for security researchers to easily automate, package, and share the AI prompts and workflows that they find effective for their work. In this blog post, I’ll share how I created auditing taskflows to find vulnerabilities in Android applications. While new models are getting better at understanding code, custom taskflow prompts let security researchers guide them—splitting research into inc

3 stories
20:00 Anthropic:Claude.dev 开发者博客(RSS)AI score 67/100

Automating eval design and hillclimbing with Claude

Evaluations provide a signal on how your app or skill is performing on specific tasks. But designing evaluations, and improving performance on them without fooling yourself, is hard. We've added guidance for both to the claude-api skill. With the skill, you can run /claude-api build-eval to build an evaluation inside your codebase, and run /claude-api hillclimb to improve your application against it, one change at a time, with a held-out set of examples to catch overfitting. In this article, we

15:33 Hacker News:AI 热帖AI score 86/100

Prompting Claude Opus 5.5

This guide covers the prompting patterns specific to Claude Opus 5.5. For the model's capabilities and API changes, see What's new in Claude Opus 5.5. For techniques that apply across all current Claude models, see Prompting best practices. Claude Opus 5.5 generates output tokens more than 30 percent faster than Claude Opus 5 and tends to finish the same task with fewer tokens. Existing Claude Opus 5 prompts should perform well without changes, and the patterns in Prompting Claude Opus 5 remain

08:00 Tomer Tunguz 博客(VC 分析)AI score 63/100

How GPU Prices Can Double While AI Gets Cheaper

GPU prices have doubled in the last six months, from $4.40 to $8.08 per GPU-hour. But AI prices are falling. How can that be? It is not a simple answer. Higher GPU costs are likely to remain while the industry races to build out new data centers. Every component of the buildout is increasing in cost, from concrete to copper to credit.1 Above all, electricity remains the limiting factor, which Oracle is experiencing : last week it invoked force majeure on its New Mexico campus after the natural-g

1 story
4 stories
23:00 GitHub BlogAI score 64/100

Improving site performance by shipping more CSS

The Primer Design System powers many of the experiences you see on GitHub today. From buttons to banners to breadcrumbs, these foundational components are required to be accessible, flexible, and performant across a wide variety of scenarios. Back in 2023, the number of components on certain pages began to explode. This led to several performance-related challenges with our existing CSS-in-JS solution: Initial page loads took longer due to styles being initialized on the client Server-side rende

3 stories
4 stories
3 stories
4 stories
1 story
4 stories
3 stories
3 stories
1 story
4 stories
3 stories
1 story
3 stories
3 stories
4 stories
16:57 Tripo(官方 X)AI score 66/100

This is what 3D Vibe Coding looks like in practice. An incredible use case combining Tripo for hig...

Tripo 转发用户案例,展示用 Images 2.5、Tripo 智能网格 P2.0、GPT-6 Astra 和 Blender MCP 以 AI 为主制作 3D 角色的完整流程。作者几乎只做视图操作,通过截图加参考图让 Astra 修正形状与纹理,并指出纹理修正在细节上仍较粗糙。作者称此前在 GPT-5.6 Sol 上反复失败的操作在 Astra 上可以直接完成。

01:52 X:Nathan Lambert (@natolambert)AI score 85/100

I never got to comment on Nvidia-HF because I was OOO, but something that struck me is how HuggingFa...

Nathan Lambert 评价 Nvidia 收购 HuggingFace,认为 HuggingFace 影响 AI 讨论方向的能力对 Nvidia 值得每年付出约 100 亿美元。他认为 Nvidia 比三大云厂商更适合做买方,HuggingFace 应摆脱盈利单位定位,去争取下一代 1 亿 AI 开发者;作者曾在 HuggingFace 工作并于去年预言过这一收购。

4 stories
08:00 OpenRouter:Announcements(RSS)AI score 64/100

Nano Banana API: Edit Images with Gemini in Code

OpenRouter 发布教程,演示通过 API 向 google/gemini-3.1-flash-image(即 Nano Banana 2)发送源图和文本指令完成图像编辑,编辑结果以 base64 形式在响应中返回。教程给出 Python 和 TypeScript 示例、提示词写法、多轮小步编辑方法,以及更换模型只需改一个字段,并介绍了 Nano Banana 系列四个成员的价格与质量差异。

00:00 Mistral AI:News(网页)AI score 63/100

Solutions Modernizing complex legacy code with AI agents. Lessons from 40,000 lines of Fortran. September 9, 2026 By Carlo Antonio Patti & Rasul Alakbarli

Mistral 帮助一家欧洲能源运营商将 40,000 行 Fortran 77 的油藏模拟器迁移到 C++,该代码库无测试套件且文档分散。团队先搭建数值一致性验证工具,用 Vibe CLI 生成调用树并派出百余个 Agent 补齐文档,最终采用由人类把关的 coder、tester、reviewer 结构化工作流逐模块迁移。

2 stories
1 story
8 stories
4 stories
00:35 Google AI:DEV 作者专属(RSS)AI score 66/100

How to Write Reliable Rubrics for LLM-as-a-Judge Evaluations

Google AI 团队发布教程,讲解如何为 LLM-as-a-Judge 评测编写可靠的布尔式评分标准,指出模糊提示会导致评估不一致和浪费 token。文中给出四条经验:问题保持原子化且互不重叠、只让评判模型评估客观事实(可用 RFC 2119 术语如 MUST 表述)、只评 prompt 中明确要求的内容、用专家标注的 golden set 校准评判模型直至与人类评分一致。

3 stories
1 story
1 story
1 story
3 stories
3 stories
1 story
Showing the latest 100 of 434 stories.