Few AI Models Profitable in 500-Day Startup Test
- Aisha Washington

- Jun 29
- 3 min read
Princeton researchers from Princeton University’s AI Lab tested fourteen language models by letting each run a simulated software company called NovaMind for 500 days. The firms started with one million dollars. Only three models finished with more capital than they began with. Lead author Dr. Elena Torres noted in the study “CEO-Bench: Evaluating Long-Horizon Decision Making in Language Models” that “most agents optimized visible output instead of survival metrics.”
The benchmark forced every agent to set prices, hire staff, manage churn, fix bugs, and launch features without human intervention. Most agents lost the entire bankroll within months. The winners were Claude Fable 5, Claude Opus 4.8, and GPT-5.5. A simple rule-based program that ignored language models entirely also beat every other entrant except the top three.
Test Setup and Rules
NovaMind sells a subscription analytics product. Agents received daily financial reports, customer tickets, and competitor moves. They had to decide on pricing tiers, marketing spend, feature priorities, and hiring. Bankruptcy ended the run. The 500-day limit matched the typical runway many early-stage founders receive before they need another round.
The three language models reached peak profits of 47.15 million, 27.8 million, and 21.3 million dollars in their best single runs. The non-model heuristic reached 15.76 million dollars. It used fixed pricing, automatic quota enforcement, and a short list of pre-approved features. Fourteen other models, including several current leaders, ended the test with zero capital.
Why Most Agents Failed
The majority of agents changed strategy every week. They raised prices after one bad sales report, then reversed course after a single complaint. They hired staff for features they later abandoned. The pattern repeated across runs: early over-expansion followed by cash burn and collapse.
Researchers noted that agents rarely checked cumulative cash flow before making new commitments. Dr. Torres observed, “Agents reacted to the latest metric instead of the trend.” That behavior produced coherent daily activity but incoherent financial results.
The Simple Heuristic Advantage
The rule-based program set three price points on day one and never changed them. It released one feature per quarter and kept a reserve equal to three months of operating costs. When revenue dipped, it cut marketing before it cut engineering. The approach produced steady growth without dramatic swings.
Its success surprised the team. The program had no ability to write marketing copy or negotiate with vendors. Yet it outlasted every model except the top three. The result suggests that consistency matters more than fluent decision text when capital is limited.
Implications for Business Strategy
Founders who deploy agents inside real companies now face a narrower set of viable options. Full autonomy remains risky unless the agent can maintain a single pricing and resource plan across hundreds of decisions. Hybrid setups that let a human lock core parameters appear safer in the near term. Beyond the simulation, the findings echo real-world patterns noted in The Verge’s coverage of AI tooling constraints in early-stage ventures.
The test also showed that CEO-Bench distinguishes long-range planning from short-term response. Models that scored well on daily coding benchmarks often failed here because they optimized for visible output rather than survival.
Limitations of the Benchmark
The simulation assumes perfect information on costs and churn. Real companies face delayed and noisy data. The 500-day horizon also ends before market cycles or regulatory changes appear. These simplifications make the gap between top and bottom models look larger than it might be in practice.
Still, the ordering of results held across multiple random seeds. The same three models and the same heuristic finished ahead of the field in nearly every repeat.
What Founders Should Watch
Teams that test agents internally should track cash runway after every pricing or hiring change, not just revenue. They should also measure how often an agent reverses its own prior decision within thirty days. High reversal rates predicted bankruptcy in the benchmark.
The next public release of CEO-Bench is scheduled for early 2027. It will add competitor pricing moves and delayed customer feedback. Those additions will test whether models can hold strategy when the immediate signal conflicts with earlier data.
How remio Fits
remio gives teams a persistent record of past pricing decisions and customer reactions. When an agent proposes a change, the record shows what happened the last time similar conditions occurred. That external memory layer reduces the reversal problem the benchmark exposed.
Founders can run the same five-hundred-day scenario inside their own data before they give an agent control of live accounts. The result is a narrower set of moves that already survived the consistency test.
https://www.remio.ai
The test makes one point clear: profitable operation over long periods still demands either exceptional model discipline or deliberate constraints on what the model is allowed to alter. Most current agents possess the first trait in limited supply.


