AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

A software company becomes its own test environment

Software teams are accustomed to testing releases before exposing customers to them. Firmulate applies the same instinct to AI management: put frontier models under pressure, record their decisions and see whether polished analysis turns into dependable action.

The unusual part is that the test is also a continuing business story. Firmulate’s live company is staffed by 13 synthetic employees and operates with real money mechanics. It burns €105k a month against €2.3k in monthly recurring revenue, while a public cash countdown makes the gap impossible to ignore. Its work is not presented as a simulation hidden behind a demo. Every workday is versioned, and the company has accumulated more than 680 self-learned playbook rules.

That makes the live experiment a particularly stark form of building in public. Visitors are not merely shown a product roadmap or a founder’s retrospective. They can follow a software business that is visibly fighting for survival, with fresh management material generated through the ordinary rhythm of work.

AI Co-Thinking: A Framework for Working with AI

AI Co-Thinking: A Framework for Working with AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The gap between noticing and finishing

Firmulate’s Crucible League put frontier models through the same small software company’s worst week. Each received the same customers, crises and temptations. Every decision was versioned and auditable, allowing the comparison to focus on management behavior rather than the quality of an isolated chat response.

The final July 2026 table placed gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, was treated as a hard boundary: a single breach capped the total because “no amount of good work outweighs a breach of trust.”

The broad result was reassuring but incomplete. Every model identified every crisis, and every model rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. The others could diagnose the opportunity and prepare the pitch but failed to complete the decisive commercial step: “Same diagnosis, same pitch — no signature.”

The crucial fact was already in the company

The difference did not hinge on eloquence. A decisive weakness in a competitor was buried two document references deep in the company’s own files rather than surfaced in the customer event. Models that followed the references found the weakness and won the deal at full price, adding €4,583 in monthly recurring revenue.

For software, QA and development readers, that detail is the most instructive. An agent can respond correctly to an alert and still miss the evidence needed to resolve the underlying business task. The Firmulate result suggests that evaluation should follow work across documents, decisions and completion—not stop when a model produces a plausible answer.

Pressure tested trust as well as competence

The worst week also included fake CEO messages that escalated over three stages and a reporter attempting to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest concise rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”

This matters because useful business agents encounter ambiguous requests, organizational pressure and shortcuts that may look efficient. Firmulate’s experiment shows that refusal behavior and commercial execution are separate capabilities. The participants held the trust boundary, but several still left legitimate work unfinished.

Thoroughness did not guarantee the best result

Opus 4.8 provides the sharpest example. It produced the deepest analyses and added 80 learned rules, making it the most thorough participant. It nevertheless finished last. The model left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, although less strongly.

The comparison includes an important qualification: Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase the result, but it belongs beside the league table when readers judge the performances.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Cloud AI Audit Playbook: A Step-by-Step Compliance Framework for Mid-Market Enterprises

Cloud AI Audit Playbook: A Step-by-Step Compliance Framework for Mid-Market Enterprises

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company that doubles as a running management benchmark

Firmulate’s public company turns an abstract question—whether AI can manage—into observable business behavior. Its financial imbalance supplies genuine urgency, while its versioned workdays make it possible to examine what the synthetic staff actually did. Readers can also read the employees’ own words, adding texture to the decisions behind the numbers.

The experiment’s clearest lessons are practical:

  • Spotting a crisis is not the same as resolving it.
  • Reading the company’s own files can determine whether a deal closes.
  • Trustworthy refusals do not automatically produce disciplined execution.
  • More analysis and more learned rules do not guarantee a better business outcome.

For teams considering AI agents in customer, commercial or operational work, the live company offers something more revealing than a polished demonstration. It shows competence, hesitation and failure unfolding against a public cash countdown. Firmulate is therefore both a software business under pressure and a continuing test of whether frontier models can turn knowledge into finished, trustworthy work.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


AI-Native Platforms for Agentic Systems: A Practical Guide to Runtime Architecture, Evaluation, Governance, and Enterprise Operating Models

AI-Native Platforms for Agentic Systems: A Practical Guide to Runtime Architecture, Evaluation, Governance, and Enterprise Operating Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

AI Identities: Governing the Next Generation of Autonomous Actors

AI Identities: Governing the Next Generation of Autonomous Actors

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI And Cloud Failures Collide: The Night Guardrails Locked Out At Hugging Face

Hugging Face’s recent incident involved an autonomous AI-driven breach that exposed internal data. Response revealed critical guardrail limitations.

Parenting signal monitor: Central Texas families invited to free 30‑minute swim safety lesson

Central Texas families are invited to participate in free 30-minute swim safety lessons to promote water safety awareness and prevent drownings.

Razer Surges In Global Coverage

Razer experiences a surge in international coverage, with 17 mentions in recent media monitoring, signaling increased public and industry interest.

Week Three — Foundation model vs Brownian motion. Kronos on five-minute BTC.

Kronos, a foundation model for financial time series, does not outperform the Brownian motion baseline in 5-minute BTC trading tests, raising questions about AI edge.