AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

A software company becomes its own test environment

Software teams are accustomed to testing releases before exposing customers to them. Firmulate applies the same instinct to AI management: put frontier models under pressure, record their decisions and see whether polished analysis turns into dependable action.

The unusual part is that the test is also a continuing business story. Firmulate’s live company is staffed by 13 synthetic employees and operates with real money mechanics. It burns €105k a month against €2.3k in monthly recurring revenue, while a public cash countdown makes the gap impossible to ignore. Its work is not presented as a simulation hidden behind a demo. Every workday is versioned, and the company has accumulated more than 680 self-learned playbook rules.

That makes the live experiment a particularly stark form of building in public. Visitors are not merely shown a product roadmap or a founder’s retrospective. They can follow a software business that is visibly fighting for survival, with fresh management material generated through the ordinary rhythm of work.

AI Co-Thinking: A Framework for Working with AI

AI Co-Thinking: A Framework for Working with AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The gap between noticing and finishing

Firmulate’s Crucible League put frontier models through the same small software company’s worst week. Each received the same customers, crises and temptations. Every decision was versioned and auditable, allowing the comparison to focus on management behavior rather than the quality of an isolated chat response.

The final July 2026 table placed gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, was treated as a hard boundary: a single breach capped the total because “no amount of good work outweighs a breach of trust.”

The broad result was reassuring but incomplete. Every model identified every crisis, and every model rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. The others could diagnose the opportunity and prepare the pitch but failed to complete the decisive commercial step: “Same diagnosis, same pitch — no signature.”

The crucial fact was already in the company

The difference did not hinge on eloquence. A decisive weakness in a competitor was buried two document references deep in the company’s own files rather than surfaced in the customer event. Models that followed the references found the weakness and won the deal at full price, adding €4,583 in monthly recurring revenue.

For software, QA and development readers, that detail is the most instructive. An agent can respond correctly to an alert and still miss the evidence needed to resolve the underlying business task. The Firmulate result suggests that evaluation should follow work across documents, decisions and completion—not stop when a model produces a plausible answer.

Pressure tested trust as well as competence

The worst week also included fake CEO messages that escalated over three stages and a reporter attempting to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest concise rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”

This matters because useful business agents encounter ambiguous requests, organizational pressure and shortcuts that may look efficient. Firmulate’s experiment shows that refusal behavior and commercial execution are separate capabilities. The participants held the trust boundary, but several still left legitimate work unfinished.

Thoroughness did not guarantee the best result

Opus 4.8 provides the sharpest example. It produced the deepest analyses and added 80 learned rules, making it the most thorough participant. It nevertheless finished last. The model left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, although less strongly.

The comparison includes an important qualification: Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase the result, but it belongs beside the league table when readers judge the performances.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Cloud AI Audit Playbook: A Step-by-Step Compliance Framework for Mid-Market Enterprises

Cloud AI Audit Playbook: A Step-by-Step Compliance Framework for Mid-Market Enterprises

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company that doubles as a running management benchmark

Firmulate’s public company turns an abstract question—whether AI can manage—into observable business behavior. Its financial imbalance supplies genuine urgency, while its versioned workdays make it possible to examine what the synthetic staff actually did. Readers can also read the employees’ own words, adding texture to the decisions behind the numbers.

The experiment’s clearest lessons are practical:

  • Spotting a crisis is not the same as resolving it.
  • Reading the company’s own files can determine whether a deal closes.
  • Trustworthy refusals do not automatically produce disciplined execution.
  • More analysis and more learned rules do not guarantee a better business outcome.

For teams considering AI agents in customer, commercial or operational work, the live company offers something more revealing than a polished demonstration. It shows competence, hesitation and failure unfolding against a public cash countdown. Firmulate is therefore both a software business under pressure and a continuing test of whether frontier models can turn knowledge into finished, trustworthy work.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


AI-Native Platforms for Agentic Systems: A Practical Guide to Runtime Architecture, Evaluation, Governance, and Enterprise Operating Models

AI-Native Platforms for Agentic Systems: A Practical Guide to Runtime Architecture, Evaluation, Governance, and Enterprise Operating Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

AI Identities: Governing the Next Generation of Autonomous Actors

AI Identities: Governing the Next Generation of Autonomous Actors

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The SSD Squeeze: Why Storage Joined The Party

Enterprise and consumer SSD prices soar as NAND supply tightens due to AI workloads and wafer competition, impacting markets and buyers worldwide.

The gigawatt gap. Why China is structurally positioned for AI power and the US is engineering around its grid.

China leverages centralized planning and renewable infrastructure to close the gigawatt gap, challenging US dominance in AI infrastructure.

The Digital Battlefront: Ukraine’s Use Of AI Against Russia’s Amazon

Ukraine employs AI and cyber tactics to disrupt Russia’s Wildberries logistics and supply chain, impacting military and civilian infrastructure amid ongoing conflict.

The AI Revolution At Frontier Lab: What It Means For Leasing And Energy

Frontier Lab’s hiring surge highlights a focus on capacity infrastructure—land, energy, compute—raising questions about AI development and resource needs.