GPTProto

Deployment & Ops

555 picksNewest firstLatest article Oct 1, 2026, 8:00 AM
7 stories
08:00 OpenRouter:Announcements(RSS)AI score 68/100

Confidence Thresholds for Model Escalation Routing

Sending every request to your strongest model gets good answers at the highest price, because you pay frontier rates for requests a cheaper model would have answered correctly. Fixed routing rules, such as picking the model by keyword or task type, cost less but need rewriting as your traffic changes. Confidence-based escalation sits between the two. You ask the model to score its own answer, then route on that score. Answers with a high score stay on a cheap model. Answers with a low score go t

08:00 Modal 官方工程博客(RSS)AI score 61/100

VM Sandboxes: Full computers for agents

Today we’re making VM Sandboxes generally available on Modal, built for those who need to give their agents the power of a full computer.With one flag, you’ll get a fully capable Linux VM with all the niceties that you expect from a traditional modal.Sandbox, and it Just Works™. This brings the same APIs, modal.Images, sub-second cold-starts, and CPU/memory bursting capabilities as previously, all whilst supporting the hundreds of thousands of concurrent Sandboxes that our users are accustomed t

08:00 OpenRouter:Announcements(RSS)AI score 67/100

How to Gate Pull Requests on LLM Evals in CI

Changing one line in a support agent’s system prompt can ship an agent that tells customers the refund window is 30 days when your policy says 14. Nothing in a normal CI pipeline checks what the model says, so the build passes and the first person to see the wrong answer is a customer.Gating a pull request on a fixed eval set works the same way as gating on a failing unit test. You keep test cases in the repository, run them when a prompt changes, and block the merge when too many fail.In this g

08:00 Modal 官方工程博客(RSS)AI score 63/100

Runtime Roundup: VM Sandboxes, Multi-node clusters, and more

Modal just hosted our inaugural conference, Runtime. Here are a few of the highlights that we announced.VM SandboxesVM Sandboxes give your agent access to a full Linux computer. Agents increasingly want to live inside something that looks like a real machine: running Docker stacks, local databases and dev servers, graphical environments and mobile simulators, and even monkeying around with the Linux Kernel itself.VM Sandboxes are already in use at customers like Linear, Legora, and Snorkel power

08:00 OpenRouter:Announcements(RSS)AI score 67/100

Cost vs. Quality Tradeoff Framework for Agent Models

Picking a model for an agent by leaderboard rank pays frontier prices for tasks that a cheaper model may handle at the same accuracy. The question to answer is not which model scores highest, but which model is the cheapest one that is good enough for the task in front of you. This guide is a three-step framework for making that call. You set the quality bar the task needs, measure cost per quality point on your own examples, and pick the cheapest model that clears the bar with margin. Tl;dr Set

04:01 Google DeepMind:Blog(RSS)AI score 78/100

Gemini 4 Argon: our next era of frontier intelligence

Sep 30, 2026 | Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. In this article Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program. Built to sustain deep reasoning across complex, long-horizon workflows, Argon is fundamentally changing the way we work and build at Go

3 stories
08:00 OpenRouter:Announcements(RSS)AI score 65/100

Building a Golden Eval Dataset from Production Traffic

You update a prompt, or a provider rolls out a new checkpoint under the same model ID, and something in production regresses. The last week of user complaints looks slightly different from the week before. A public benchmark like MMLU won’t catch that. It measures general capability across academic subjects, not how the model handles your product’s traffic. A golden eval dataset closes that gap. It’s a curated collection of production inputs paired with reviewed expected outputs, versioned in Gi

08:00 OpenRouter:Announcements(RSS)AI score 69/100

AI Agent Regression Testing After a Prompt or Model Change

A ~author/family-latest alias always resolves to the newest concrete model in a family. That’s convenient in production and a problem in a regression test, because the model can change between runs without any change in your repository. Our latest model resolution docs describe the mechanism and recommend a concrete model slug when you need a fixed version for reproducibility. This guide covers the locked case set and the behavioral contract, then the model-swap case in detail. Tl;dr Regression

01:20 OpenAI:官网动态(RSS · 排除企业/客户案例)AI score 86/100

Introducing GPT-6.1 Sol

Near-Astra intelligence for a fifth of the price We’re introducing GPT‑6.1 Sol, an upgrade to GPT‑6 Sol that nearly matches GPT‑6 Astra’s intelligence on agentic coding, computer use, and professional work at one-fifth of Astra’s standard input and output token prices. Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing—giving developers more room to build and run capable agents that reuse context across request

2 stories
3 stories
20:00 Anthropic:Claude.dev 开发者博客(RSS)AI score 67/100

Automating eval design and hillclimbing with Claude

Evaluations provide a signal on how your app or skill is performing on specific tasks. But designing evaluations, and improving performance on them without fooling yourself, is hard. We've added guidance for both to the claude-api skill. With the skill, you can run /claude-api build-eval to build an evaluation inside your codebase, and run /claude-api hillclimb to improve your application against it, one change at a time, with a held-out set of examples to catch overfitting. In this article, we

15:33 Hacker News:AI 热帖AI score 86/100

Prompting Claude Opus 5.5

This guide covers the prompting patterns specific to Claude Opus 5.5. For the model's capabilities and API changes, see What's new in Claude Opus 5.5. For techniques that apply across all current Claude models, see Prompting best practices. Claude Opus 5.5 generates output tokens more than 30 percent faster than Claude Opus 5 and tends to finish the same task with fewer tokens. Existing Claude Opus 5 prompts should perform well without changes, and the patterns in Prompting Claude Opus 5 remain

08:00 Tomer Tunguz 博客(VC 分析)AI score 63/100

How GPU Prices Can Double While AI Gets Cheaper

GPU prices have doubled in the last six months, from $4.40 to $8.08 per GPU-hour. But AI prices are falling. How can that be? It is not a simple answer. Higher GPU costs are likely to remain while the industry races to build out new data centers. Every component of the buildout is increasing in cost, from concrete to copper to credit.1 Above all, electricity remains the limiting factor, which Oracle is experiencing : last week it invoked force majeure on its New Mexico campus after the natural-g

2 stories
4 stories
04:12 X:Eric Zakariasson (@ericzakariasson)AI score 77/100

here's a prompt to improve your agent harness based on what we've learned at cursor. enjoy # Improv...

一份公开提示词用于优化 LLM Agent Harness,目标是在不降低任务质量的前提下降低每任务的价格加权 token 成本。某团队一轮改动(提示词精简、工具卸载、缓存布局、稀疏行号、子智能体调优)将整体 token 成本降低约 7%,且质量无损。提示词强调按任务而非按请求计量,并建议先映射 harness、测量基线,再按优先级实施改动。

4 stories
8 stories
2 stories
1 story
2 stories
1 story
2 stories
2 stories
1 story
2 stories
3 stories
5 stories
2 stories
6 stories
4 stories
08:00 xAI:News(网页)AI score 62/100

Setting Grok Bot loose on procurement We gave Grok Bot access to vendor spend, contracts, and usage data. It found more than $100,000 in direct savings. Sep 4, 2026

xAI 给 Grok Bot 开放供应商支出、合同和使用数据,搭建了名为 Haggle Bot 的采购智能体,已找出超过 10 万美元直接节省。SaaS 审计中发现一个产品有 43 个付费席位 90 天无活跃,节省 $14,220,另一个产品每年有 $85,662 未使用 SKU;办公用品比价一次将 $14,629 的订单压到 $6,143,降幅 58%。

4 stories
00:35 Google AI:DEV 作者专属(RSS)AI score 66/100

How to Write Reliable Rubrics for LLM-as-a-Judge Evaluations

Google AI 团队发布教程,讲解如何为 LLM-as-a-Judge 评测编写可靠的布尔式评分标准,指出模糊提示会导致评估不一致和浪费 token。文中给出四条经验:问题保持原子化且互不重叠、只让评判模型评估客观事实(可用 RFC 2119 术语如 MUST 表述)、只评 prompt 中明确要求的内容、用专家标注的 golden set 校准评判模型直至与人类评分一致。

5 stories
00:00 Claude:Blog(网页)AI score 79/100

A guide to the anatomy of effective commerce agents

Anthropic 发布构建电商智能体的指南,总结架构、延迟与成本优化、生产运维三部分实践,涉及购物和商家两类智能体,企业客户部署后出现购物车变大和卖家运营效率提升。核心建议包括用单个智能体加技能而非子智能体、将 UI 组件做成工具、用提示词缓存实现 90-99% 命中率、在 harness 中强制安全规则,并开源参考实现 anthropics/commerce-agents。

1 story
2 stories
2 stories
2 stories
4 stories
08:00 Claude Platform:开发者版本说明(RSS)AI score 63/100

Claude Platform release notes — August 27, 2026

Claude Console 现已支持创建个人密钥和服务账号密钥,它们以关联账户身份运行并继承相同权限,账户从组织移除后密钥即失效。组织管理员可借此更轻松追踪各账户用量并确保密钥使用合规。这些 API 密钥可限定到特定工作区,也可用于管理端点及账户可访问的任何工作区,工作区 API 密钥仍作为旧版选项保留支持。

00:00 LMSYS:Blog(Chatbot Arena 团队)AI score 65/100

Blog MiniMax-H3 on 8×H200: 1.95× Lossless, Up to 6.24× at 0.76–0.91 SSIM We benchmarked MiniMax-H3 video generation on 8× NVIDIA H200 with SGLang Diffusion, holding prompts, seeds, resolution, frame rate, and denoising steps fixed across six workloads. - SGLang's dense, l... SGLang Diffusion Team, Cache-DiT Team, NVIDIA, Ant Group

SGLang Diffusion 团队在 8×NVIDIA H200 上对 MiniMax-H3 视频生成进行基准测试,其密集无损路径较 Diffusers 快 1.85–1.95×,无近似损失。

4 stories
4 stories
08:00 OpenRouter:Announcements(RSS)AI score 68/100

How to Choose the Best AI Model (Live, in Your Editor)

OpenRouter 提出一套模型选型框架:先定义任务,从实时用量和第三方基准中筛选候选,再对比各提供商的定价与延迟,最后用自有提示词测试。判断标准是“每完成任务的成本”而非“每 token 成本”。其 MCP 服务器可直接在 Claude Code、Cursor 等编辑器中查询实时排名、价格和基准,并用 openrouter/auto-beta 按请求路由。

02:02 Meta Engineering Blog(RSS)AI score 62/100

MetaRoCE: A New RDMA Transport Built for AI-Scale Ethernet

Meta 设计并开源了 MetaRoCE,一个专为 AI 工作负载在通用以太网上打造的 RDMA 传输协议,已通过 Open Compute Project(OCP)发布规范、参考软件实现和合规测试套件。该协议将智能移至端点,原生支持乱序交付、多路径、无损容忍和双向拥塞控制,无需 PFC,可在百万 GPU 规模下提供高吞吐、低尾延迟。现有 RDMA Verbs API 和软件栈无需修改即可运行。

5 stories
1 story
Showing the latest 100 of 555 stories.