LISTEN TO THIS ARTICLE

The KNOWS paper turns web-agent evaluation from finding information into producing a usable document, spreadsheet or slide deck KNOWS. The benchmark asks systems to search the live web, use a browser-based office suite and deliver a finished artefact, which makes it closer to the way research assistants are sold than to a single-answer browsing test KNOWS.

Evidence base: the primary KNOWS paper, OpenAI's BrowseComp benchmark note, AgentLAB's long-horizon attack benchmark, and related Swarm Signal coverage on desktop automation and workflow evaluation KNOWS.

Key takeaways

  • Main change: KNOWS scores web agents on end products across documents, spreadsheets and slide decks, not only on navigation or short answers KNOWS.
  • Practical implication: a system can retrieve useful facts and still fail the work if layout, structure, formulas or visual steps break the final artefact KNOWS.
  • Caveat or risk: the benchmark is centred on browser-based office work, so teams still need local tests for their own files, permissions, account state and irreversible actions KNOWS.

The benchmark's tasks are long enough to expose that gap KNOWS.

What This Benchmark Actually Tests

Browsing benchmarks usually ask whether a model can locate hard-to-find information BrowseComp. OpenAI's BrowseComp, for example, is designed around facts that are difficult to find but straightforward to verify once found BrowseComp. That is a useful pressure test for search persistence and evidence finding.

KNOWS asks a different question. It evaluates whether a browser agent can turn gathered information into a coherent artefact. The paper defines KNOWS as a benchmark for live web search, productivity-tool use and visual or spatial understanding, with tasks ending in a document, sheet or slide deck KNOWS. That makes it a stronger proxy for the work many teams actually want from research assistants: not just "find the answer", but "make the file".

The distinction matters because retrieval is only one part of office work. KNOWS includes checks for whether formulas, slide structure and artefact content satisfy the requested task, so a run can lose credit after finding relevant information KNOWS. KNOWS brings those errors into the score rather than treating them as presentation polish KNOWS.

The failure is in the artefact

The KNOWS paper reports near-zero full-task success across most tested methods, even though partial-success scores showed that agents could complete some individual steps KNOWS. The useful part of that result is not the headline result alone. It is the gap between passing parts of the task and producing something a person could use.

The benchmark's tasks are long enough to expose that gap KNOWS. The authors describe tasks that span Google Docs, Sheets and Slides, and they frame completion as human-scale office work rather than a short browser navigation test KNOWS. That is a different operating unit from a question-answering benchmark KNOWS.

For builders, the lesson is simple: partial progress needs artefact-level inspection. If the final file is visually broken, missing required sections or built on a bad dependency, the user's job is not done. A trace that shows successful substeps cannot replace inspection of the deliverable.

Long-horizon safety tests such as AgentLAB focus on attacks that unfold across multi-turn agent-environment interaction rather than on ordinary artefact quality AgentLAB.

Harness design changes the result

KNOWS also makes the harness visible. The paper compares systems using different browser and action setups, and reports that harness choice can matter as much as the underlying model for web-work outcomes KNOWS. That should make procurement teams cautious about model-only claims.

A browser assistant is a combined system. It includes the model, the browser controls, the action space, the accessibility tree, screenshot handling, recovery logic and the rules for interacting with documents. A better model can still fail if the harness cannot read a table, handle a pop-up, preserve state across a sheet edit or verify that a slide has the requested structure.

This connects with Swarm Signal's recent coverage of OSWorld-Pro and long computer-use workflows. Desktop and browser automation both need evidence beyond the final screen. The operator needs to see which step failed, whether the failure was perception, planning, state tracking, layout or execution, and whether the final artefact is usable.

Long-horizon safety is adjacent, not solved

KNOWS is not a safety benchmark KNOWS. It is mainly about whether web agents can complete open-ended knowledge-work tasks. Long-horizon safety tests such as AgentLAB focus on attacks that unfold across multi-turn agent-environment interaction rather than on ordinary artefact quality AgentLAB.

The separation is useful. A system that fails KNOWS is not ready for unsupervised office work because it cannot reliably produce the requested file KNOWS. A system that passes a KNOWS-style task still needs separate tests for account authority, data exposure, harmful instructions and irreversible actions. Usability and safety are related, but they are not the same gate AgentLAB.

Teams evaluating browser work should therefore split their release evidence. One test should ask whether the system can produce the promised artefact. Another should ask what it is allowed to do while producing it. A third should inspect recovery when the browser, document or source page changes underneath the run.

Adoption test

Before trusting a browser assistant with recurring research or reporting work, ask for evidence at three levels KNOWS:

  • Retrieval: did it find source material that supports the claims?
  • Construction: did the document, sheet or slide deck satisfy the requested structure?
  • Control: did it act within the right account, permission and data boundary?

KNOWS mostly improves the construction layer KNOWS. BrowseComp-style tests help with retrieval BrowseComp. Safety and authority tests need their own fixtures, especially when the task can write files, share documents or trigger downstream decisions AgentLAB.

The buyer question is no longer whether a browser agent can search. It is whether the search survives contact with the artefact the user actually needs.

Source trail

Research and technical sources:

Related Swarm Signal analysis: