Era’s Synthetic Companies Expose the Agent Lookup Gap
Executive Summary
Enterprise agents can look dependable on simple lookups while failing at the joins and calculations that make business questions useful. Eon’s Era benchmark makes that gap concrete: across nine models, reported mean accuracy was 92.6% on simple filtering and 3.7% on multi-hop questions. Its more consequential contribution is a test environment where those harder questions have computable answers because every application describes the same fictional company.
This is a narrowly scoped, developer-reported result—not a verdict on enterprise agents generally. The comparison covers one company and 33 questions under one tool-use setup. The strongest lesson is about evaluation design: plausible synthetic documents are not enough to establish whether an agent understands a business.
Generate the Company Before the Documents
In “LLMs Suck at Generating Synthetic Data”, Era engineer Benji Gruenbaum describes a familiar trap: individually convincing records can collectively depict an impossible company. Employees appear before their hire dates; deals and support tickets reference entities that do not agree across systems. The critique concerns relational consistency, not whether models can write believable prose.
Era reverses the process. A seeded entity graph establishes employees, customers, ownership, and business history first. Product simulators then project that shared graph into vendor-specific records. A call in Gong and its related Salesforce activity describe the same event; a support ticket belongs to a customer that actually exists. The paper describes 66 product simulators, although the reported model comparison uses only Salesforce, Zendesk, and Gong.
The distinction matters because a cross-system question needs both consistent records and a known answer. Era computes expected answers from the completed data, while read-only access prevents an agent from changing the facts it is being graded against.
Delayed discovery: the underlying Era by Eon benchmark paper was submitted September 9. This is an older research result brought into focus by the field note, not a newly published benchmark today; the field note’s exact publication date is unconfirmed.
The Gap Is Larger Than a Leaderboard
The authors ran each of nine models on the same 33 questions three times: 891 answers in total. Besides the filtering and multi-hop extremes, mean accuracy was 39.6% for cross-system questions, 11.1% for ranking/order-sensitive questions, and 7.4% for temporal questions. These capability categories should not be treated as a partition to average together.
One revealing task asks an agent to identify the customer with the largest total value of open sales opportunities, then calculate that customer’s recorded call time. It appeared among the five hardest questions for seven models. No model correctly answered the median ticket-resolution-time question.
Those results measure a particular interaction design as well as model capability. Agents discover available systems and tools through three meta-tools; they receive budgets of 25 turns and 40 tool calls, with responses capped at 8,000 characters. Grading gives no partial credit. Poor results do not isolate whether the failure came from discovery, incomplete retrieval, aggregation, arithmetic, or exhausted budgets.
Nor does the study establish a reliable model ranking: only three of 36 pairwise comparisons remained supported after statistical correction. Repeating questions reduces generation noise; it does not turn 33 question shapes into a broad enterprise workload.
Why the Test Environment Matters
The paper’s most useful methodological detail is an audit of answer reachability. One early question had a correct join in the underlying data but exposed no usable join key through the product interfaces. The audit caught it. Ground truth alone cannot make a benchmark fair if an agent cannot retrieve the facts needed to reproduce it.
Realism remains another limitation. The authors improved their own realism score and eliminated flags from their own synthetic-artifact detector, but that is not independent proof of production fidelity. Some reference targets are estimates, and templated text lacks the diversity of human writing. The reported comparison also excludes document tasks and internal-database questions.
This reinforces the recurring distinction between generating an impressive artifact and establishing dependable performance. For enterprise agents, the next meaningful test is not another polished lookup demo: it is whether the system can exhaustively retrieve, join, and calculate over a coherent business history—and admit when the answer is absent.
Further Reading
- Gruenbaum’s field note: the practical case for generating connected business data rather than isolated documents.
- The Era benchmark paper: worth preserving for interface-level answer audits, exact grading, and explicit limits on the model comparison.

