Give GPT-5 mini the same web task and change only its interface. In ComponentBench, accessibility-tree observations produced an 83.1%srcComponentBench makes the interface part of the test, Results Table 1 task-success rate; coordinate-only Pixel control produced 48.9%srcComponentBench makes the interface part of the test, Results Table 1, a 34.2srcComponentBench makes the interface part of the test, Results Table 1-point swing inside the shared harness. In FinRCA-Bench, Typed Provenance Graph Retrieval (TPGR) and Dense retrieval-augmented generation (Dense RAG) used the same downstream language model, yet exact accuracy was 72.44%srcFinancial AI Can Get the Cause Right With Incomplete Evidence, Sections 7.2 and 7.4; Table 6 versus 2.05%srcFinancial AI Can Get the Cause Right With Incomplete Evidence, Sections 7.2 and 7.4; Table 6; TPGR achieved strict full-contract coverage in 80 of 437 retrieval-evaluable cases. The weights did not move. The route to the answer did.
The week’s through-line is that capability is increasingly determined by the interface around a model: its retrieval grammar, observation space, tool knowledge, and evidence contract. Those surfaces are now productive levers; they can give a fixed model access to records, widgets, instruments, or tools. They can also become sources of false confidence when a route, transcript, or label is mistaken for a completed and defensible result. As model routers, Model Context Protocol (MCP) servers, agent runtimes, and endpoint controls ship, the interface deserves to be treated as part of the model’s claim.
ComponentBench makes the interface effect unusually clean. Its 2,910srcComponentBench makes the interface part of the test, Abstract and Methods short tasks span 97srcComponentBench makes the interface part of the test, Abstract and Methods canonical component types, and the shared harness compares accessibility-tree (AX-tree), Set-of-Marks (SoM), and Pixel conditions. GPT-5 mini’s 83.1% AX-tree and 48.9% Pixel results are a 34.2-point change in the measured system. The wider table resists a fixed hierarchy: GPT-5.4srcComponentBench makes the interface part of the test, Results Table 1 reached 83.8%srcComponentBench makes the interface part of the test, Results Table 1 with Pixel against 81.5%srcComponentBench makes the interface part of the test, Results Table 1 with AX-tree, while GPT-5.4 mini went from 79.1%srcComponentBench makes the interface part of the test, Results Table 1 with AX-tree to 77.1%srcComponentBench makes the interface part of the test, Results Table 1 with Pixel. Model and interface interact. FinRCA applies the same logic to records. It freezes gpt-5.6srcFinancial AI Can Get the Cause Right With Incomplete Evidence, Sections 5.6 and 7.4-sol, the prompt, taxonomy, structured output schema, output limit, and retry policy, then changes retrieval. Dense RAG recovered 0.83%srcFinancial AI Can Get the Cause Right With Incomplete Evidence, Sections 7.2 and 7.4; Table 6 macro required-record recall and reached 2.05% exact accuracy; TPGR reached 77.70%srcFinancial AI Can Get the Cause Right With Incomplete Evidence, Sections 7.2 and 7.4; Table 6 and 72.44%, respectively. Its default-deny graph follows persisted transaction relationships, limits traversal to three hops, and caps the selected context at 40srcFinancial AI Can Get the Cause Right With Incomplete Evidence, Sections 5.5 and 7.4 source records. That makes provenance-aware access a plausible dominant bottleneck relative to this baseline. The benchmark also counted 254srcFinancial AI Can Get the Cause Right With Incomplete Evidence, Tables 6 and 7; Sections 7.4–7.5 correct root-cause labels with incomplete retrieval. Record access can make a diagnosis land; semantic support is a separate layer of proof.
The weights did not move. The route to the answer did.
Perplexity’s Wide ANd Deep Research (WANDR) pushes the interface question past finding a plausible page. Its 500srcWANDR finds the gap between finding facts and proving them, Opening and ceo_cfo_appointments example tasks require collections whose qualification-key trees terminate in evidence; a representative task asks for at least 70srcFinancial AI Can Get the Cause Right With Incomplete Evidence, Sections 7.2 and 7.4; Table 6 United States-based companies, with separate appointment and listing pages for each, or 140srcWANDR finds the gap between finding facts and proving them, Opening and ceo_cfo_appointments example records. Search as Code led the main run at 0.363srcWANDR finds the gap between finding facts and proving them, What we found soft F1 and 0.133srcWANDR finds the gap between finding facts and proving them, What we found hard F1, while the best hard precision and recall in the release were 0.150srcWANDR finds the gap between finding facts and proving them, What we found and 0.134srcWANDR finds the gap between finding facts and proving them, What we found. Terminal evidence-slot completion ranged from 0.979srcWANDR finds the gap between finding facts and proving them, What we found and Mean task-level raw structural completion figure to 0.994srcWANDR finds the gap between finding facts and proving them, What we found and Mean task-level raw structural completion figure across the six full runs; top-level discovery ranged from 0.611srcWANDR finds the gap between finding facts and proving them, What we found and Mean task-level raw structural completion figure to 0.951srcWANDR finds the gap between finding facts and proving them, What we found and Mean task-level raw structural completion figure. The missing work happens upstream, when the agent decides what belongs in the collection. The same diagnosis appears in physical form in Anthropic’s science demonstrations. Claude orchestrated publicly available structure-design, sequence-design, folding, and co-folding models, then Adaptyv Bio and Twist Bioscience produced and tested the candidates. In the chemistry demonstration, it turned raw nuclear magnetic resonance (NMR) and liquid chromatography–mass spectrometry (LC-MS) files into a reviewable report. The protein campaign counted 354srcClaude’s science demo is real—and narrower than the pitch, Claude’s performance on the targets binders from 1,320srcClaude’s science demo is real—and narrower than the pitch, Claude’s performance on the targets designs against 14srcClaude’s science demo is real—and narrower than the pitch, Claude’s performance on the targets of 15srcClaude’s science demo is real—and narrower than the pitch, Claude’s performance on the targets targets; the chemistry result used one routine quality-control sample. Lab testing remained the gate. Producing a candidate, citing a page, or returning a spectrum crosses an interface; completion is the next milestone.
The model may be the same. The interface is where the verdict moves.
An automatic speech recognition (ASR) study compared 11srcWhen ASR Follows the Transcript Instead of the Audio, Methods and Results opening; Figure 2 open-source recognizers on English VoxPopuli and found that the six with 5.4–5.8%srcWhen ASR Follows the Transcript Instead of the Audio, Methods and Results opening; Figure 2 word error rates (WERs) also had reference-disagreement accept-ref rates of 0.18srcWhen ASR Follows the Transcript Instead of the Audio, Methods and Results opening; Figure 2–0.30srcWhen ASR Follows the Transcript Instead of the Audio, Methods and Results opening; Figure 2; every model at 6.5%srcWhen ASR Follows the Transcript Instead of the Audio, Methods and Results opening; Figure 2 WER or higher was at 0.10srcWhen ASR Follows the Transcript Instead of the Audio, Methods and Results opening; Figure 2 or below. The probe asks what a system does when audio leaves two renderings viable. On generic voices reading the identical transcript, accept-ref fell for many models, while clones of fresh same-domain speakers often moved closer to generic voices. A low WER can carry a corpus convention; the size of any resulting WER inflation remains unmeasured. The self-improvement audit makes the same warning in a different register. Run a frozen Qwen3-8B through the same temperature-zero evaluation twice and a single-decode ledger records six apparent new solves and nine lost ones. Serializing requests removed three-quarters of the flips, yet serial evaluation still changed about 2%srcSelf-improvement needs a measured null, Introduction temperature-zero decoding and batching discussion of greedy verdicts. At k=128srcSelf-improvement needs a measured null, Introduction expansion-statistic and exact-test discussions; Conclusion E.3, a frozen comparison labeled seven of 25srcSelf-improvement needs a measured null, Introduction expansion-statistic and exact-test discussions; Conclusion E.3 American Invitational Mathematics Examination (AIME) problems as expanded. The proposed per-problem exact test made no detections on held-out frozen replicates; that all-null result provides no power estimate or minimum detectable effect. The unchanged model is the necessary baseline for a claim that a model learned.
Product releases made the architecture tangible. TrueFoundry’s MIT-licensed, vendor-neutral TrueForge runtime leaves model and infrastructure choices to developers while supplying the agent loop, tool and MCP orchestration, persistent sessions, context compaction, approvals, and sandbox-as-a-tool support. Salesforce’s Headless 360srcIndustry digest, Salesforce opens Headless 360 to external agents through MCP expansion exposes authorized capabilities through MCP; its Data 360 MCP Server offers approximately 200srcIndustry digest, Salesforce opens Headless 360 to external agents through MCP APIs, and more than 100srcIndustry digest, Salesforce opens Headless 360 to external agents through MCP reusable Agent Skills are generally available. Perplexity’s open-source Numbat adds system-level rules around agents with privileged endpoint access. The model is becoming one component inside a control surface that also owns state, permissions, and response. Ramp’s Router makes the same layer observable: it can route requests using up to three user-specified benchmarks and exposes token spend, cost, latency, and fallback data. OpenAI’s Private Safety Processing preview extends safety checks across activity in multiple conversations without retaining customer data. In deployment, the interface decides as much about trust as the model does.
Here is where the clean story frays. Interface changes often arrive as bundles. FinRCA’s contracts encode one notion of evidence sufficiency, and semantic citation support remains unadjudicated; WANDR’s benchmark and winning Search as Code system come from the same organization, with task-specific judges and no reported human-agreement study. ComponentBench’s GPT-5 mini headline is an aggregate table value without a matching interval or repeated-run estimate. The ASR comparison is cross-sectional across architectures and training mixtures, while the self-training audit is roughly 270 optimizer steps on one Qwen backbone and its all-null result has no power estimate. Better access can improve the observed answer while leaving causal reasoning unsettled; a cleaner metric can still be underpowered. The value of these studies lies in the contracts they expose. Their numbers should travel no farther than the interface and control that produced them. The interface is real; attribution is the part still on trial.
Carry an interface ledger into next week. For an agent, report the observation and action space, retrieval path, tool and schema exposure, branch completion, evidence support, and resource budget. For a learning or benchmark claim, pair the result with a frozen or no-op control, fresh voices or data, and a matched budget. The practical standard is simple: report the model and the world it was allowed to inhabit.
The model may be the same. The interface is where the verdict moves.
Grounding — claim → source
| ComponentBench compares the same task suite under accessibility-tree, Set-of-Marks, and Pixel conditions in a shared harness. | ComponentBench makes the interface part of the test, Methods and Results |
| ComponentBench contains 2,910 tasks and 97 canonical component types. | ComponentBench makes the interface part of the test, Abstract and Methods |
| GPT-5 mini succeeds on 83.1% of tasks with AX-tree observations and 48.9% with Pixel control, a 34.2-percentage-point difference. | ComponentBench makes the interface part of the test, Results Table 1 |
| GPT-5.4 reaches 83.8% with Pixel and 81.5% with AX-tree, while GPT-5.4 mini reaches 79.1% with AX-tree and 77.1% with Pixel. | ComponentBench makes the interface part of the test, Results Table 1 |
| Dense RAG and TPGR use the same downstream gpt-5.6-sol configuration, prompt, taxonomy, output schema, output limit, and retry policy. | Financial AI Can Get the Cause Right With Incomplete Evidence, Sections 5.6 and 7.4 |
| Dense RAG reaches 0.83% macro required-record recall and 2.05% exact accuracy, while TPGR reaches 77.70% macro recall and 72.44% exact accuracy. | Financial AI Can Get the Cause Right With Incomplete Evidence, Sections 7.2 and 7.4; Table 6 |
| TPGR uses a default-deny graph of persisted relationships with a three-hop limit and a 40-record selection cap. | Financial AI Can Get the Cause Right With Incomplete Evidence, Sections 5.5 and 7.4 |
| The FinRCA attribution includes 254 correct labels despite incomplete retrieval. | Financial AI Can Get the Cause Right With Incomplete Evidence, Tables 6 and 7; Sections 7.4–7.5 |
| WANDR contains 500 tasks, including a representative task requiring at least 70 United States-based companies and separate appointment and listing evidence for each, or 140 records. | WANDR finds the gap between finding facts and proving them, Opening and ceo_cfo_appointments example |
| Perplexity Search as Code leads the WANDR main run at 0.363 soft F1 and 0.133 hard F1, while the best hard precision and recall are 0.150 and 0.134. | WANDR finds the gap between finding facts and proving them, What we found |
| Terminal evidence-slot completion ranges from 0.979 to 0.994 across six full runs, while top-level discovery ranges from 0.611 to 0.951. | WANDR finds the gap between finding facts and proving them, What we found and Mean task-level raw structural completion figure |
| Claude orchestrates structure-design, sequence-design, folding, and co-folding models, while Adaptyv Bio and Twist Bioscience produce and test the designs. | Claude’s science demo is real—and narrower than the pitch, The campaign |
| Claude converts raw NMR and LC-MS files into calibrated spectra, peak tables, mass and ultraviolet spectra, and purity outputs. | Claude’s science demo is real—and narrower than the pitch, Claude runs the analytical chemistry workflow |
| The protein campaign reports 354 binders from 1,320 designs against 14 of 15 targets. | Claude’s science demo is real—and narrower than the pitch, Claude’s performance on the targets |
| The chemistry demonstration uses one routine quality-control sample. | Claude’s science demo is real—and narrower than the pitch, Claude runs the analytical chemistry workflow |
| The ASR study evaluates 11 open-source models on English VoxPopuli, and its six models with 5.4%–5.8% WER have reference-disagreement accept-ref rates of 0.18–0.30, while models at 6.5% WER or higher are at 0.10 or below. | When ASR Follows the Transcript Instead of the Audio, Methods and Results opening; Figure 2 |
| Generic voices reading the identical transcript reduce accept-ref for many models, while fresh same-domain speaker clones often move closer to the generic condition. | When ASR Follows the Transcript Instead of the Audio, Results voice-clone analysis; Figure 4 |
| A frozen Qwen3-8B evaluated twice with temperature-zero decoding produces six apparent new solves and nine lost ones; serialization removes three-quarters of the flips, while serial evaluation still changes about 2% of greedy verdicts. | Self-improvement needs a measured null, Introduction temperature-zero decoding and batching discussion |
| At k=128, a frozen comparison labels seven of 25 AIME problems as expanded, while the proposed exact test makes no detections on held-out frozen replicates. | Self-improvement needs a measured null, Introduction expansion-statistic and exact-test discussions; Conclusion E.3 |
| The all-null held-out result does not establish test power or a minimum detectable effect. | Self-improvement needs a measured null, Introduction and Abstract statistical interpretation |
| TrueForge is an MIT-licensed, vendor-neutral runtime that supplies agent loops, tool and MCP orchestration, persistent sessions, context compaction, approvals, and sandbox-as-a-tool support while leaving model and infrastructure choices to developers. | Industry digest, TrueFoundry open-sources a vendor-neutral agent runtime |
| Salesforce’s Data 360 MCP Server exposes approximately 200 APIs, and more than 100 reusable Agent Skills are generally available. | Industry digest, Salesforce opens Headless 360 to external agents through MCP |
| Perplexity’s Numbat is an open-source security suite for agent harnesses on enterprise endpoints, including privileged agents. | Industry digest, Perplexity open-sources Numbat for endpoint agent security |
| Ramp’s Router can route requests using up to three user-specified benchmarks and exposes token spend, cost, latency, and fallback data. | Industry digest, Ramp launches Router, a multi-provider model-routing API |
| OpenAI’s Private Safety Processing preview monitors activity across multiple conversations without retaining customer data. | Industry digest, OpenAI previews Private Safety Processing alongside Zero Data Retention |
| FinRCA’s evidence contracts encode one notion of sufficiency, lack a minimal-sufficient-evidence ablation, and leave semantic citation support unadjudicated. | Financial AI Can Get the Cause Right With Incomplete Evidence, Sections 3.2, 6.3, 8.4, and 9.5 |
| WANDR’s authors control both the benchmark and the winning Search as Code system, and the evaluation reports no human-agreement study or judge-error analysis. | WANDR finds the gap between finding facts and proving them, forensic_context.red_flags and forensic_context.ledger |
| The GPT-5 mini ComponentBench headline is an aggregate table percentage without a corresponding interval or repeated-run estimate. | ComponentBench makes the interface part of the test, Results statistical comparison discussion |
| The ASR comparison is cross-sectional across architectures, parameter scales, and training mixtures, while the self-training experiment uses roughly 270 optimizer steps on one Qwen backbone family. | When ASR Follows the Transcript Instead of the Audio, Figure 2 and Conclusion; Self-improvement needs a measured null, Conclusion D.1 |
| The studies recommend reporting no-op or frozen controls, fresh-speaker or fresh-data probes, evidence contracts, and matched resource budgets alongside headline scores. | Self-improvement needs a measured null, Conclusion recommendations; When ASR Follows the Transcript Instead of the Audio, practical contribution; WANDR finds the gap between finding facts and proving them, What’s next |
| MidTool-Mix is an additional 20.3B-token stage before supervised fine-tuning and reinforcement learning. | MidTool shows a tool-use gain, not yet a mid-training effect, Abstract and Baselines and Training Setup |
| GRIP uses a four-dimensional noisy evidence channel and bundles it with premise selection, span extraction, natural-language-inference filtering, and compression. | GRIP’s RAG gains outrun its causal story, Introduction, Methods, and Results |
| SparsePR reports roughly 22%–26% realized executed-pair density with aggregate quality close to dense attention, while online work accounting lacks a wall-clock breakdown. | SparsePR reconstructs skipped attention, with a speedup still to audit, Results Table 1 and Implementation details |