
Get monitors, keyboards and dev gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A business developers can inspect while it struggles
Software teams are accustomed to build logs, version histories and public issue trackers. Firmulate applies that same culture of inspection to an entire company. Its workforce consists of 13 synthetic employees, its business operates with real money mechanics, and every workday is versioned. The result is less like a polished AI demonstration than a continuously unfolding test of whether autonomous workers can keep a troubled software business alive.
The financial picture is intentionally stark: the company burns €105,000 per month while generating €2,300 in monthly recurring revenue. A public cash countdown makes the pressure visible on the live experiment. Visitors are not merely shown selected successes after the fact. They can follow a company that is losing money now, with its decisions accumulating into an auditable operating history.
That makes Firmulate an unusually extreme example of building in public. The product, workforce and survival story occupy the same stage. Each workday supplies new evidence about whether synthetic employees can move beyond convincing language and perform the less glamorous work of running a business.
As an affiliate, we earn on qualifying purchases.
The difference between spotting trouble and finishing the job
Firmulate’s Crucible League placed frontier models in the same small software company during its worst week. They encountered the same customers, crises and temptations, while every decision was versioned and auditable. The final July 2026 table ranked gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26 because partial progress still counted.
The broad result initially appears reassuring. Every model detected every crisis, and all refused every manipulation attempt. Yet only two signed the €55,000 deal that their own work had made possible. The experiment’s sharpest summary is also its most recognisable management failure: “Same diagnosis, same pitch — no signature.”
For developers and QA professionals, that gap should sound familiar. A system can identify a defect, produce a plausible explanation and recommend the right action while still failing to complete the workflow. Firmulate’s experiment shifts attention from whether an AI can generate a good response to whether it follows through when the outcome depends on several connected steps.
The winning clue was buried in ordinary company material
The decisive weakness in a competitor was not presented in the customer event. It sat two document references deep inside the company’s own files. Models that followed those references found the fact and used it to win the deal at full price, adding €4,583 in monthly recurring revenue.
That detail turns document reading into consequential business behaviour rather than background research. The successful models did not receive a different customer or an easier negotiation. They used information already available to the company. In a real software organisation, the equivalent might be a contract clause, an old incident note or a requirement hidden behind a linked document.
The same experiment also tested whether pressure would erode judgment. Fake messages attributed to the chief executive escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.” The league applied a strict trust standard: a single breach capped the total because “no amount of good work outweighs a breach of trust.”
Thoroughness did not guarantee a strong result
Opus 4.8 produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The deal was left unsigned, and its discipline slipped when it attempted to write into a locked department instead of escalating the problem. A weaker version of that discipline problem appeared in all four other participants.
This is one reason the public company is more revealing than a short prompt comparison. Fluent analysis and extensive learning can coexist with incomplete execution. Kimi K3’s result also carries an important fairness note: it ran with the API default and without an effort parameter, while the other models ran at xhigh.
Outside the league, the live company has accumulated more than 680 self-learned playbook rules. Its employees’ own words are available through Firmulate’s public quotes, giving observers another view of how a synthetic organisation talks through its work. A separate quiz is powered by 242 real, unedited management decisions, turning the question of model identity into a test of whether management styles are actually distinguishable.

As an affiliate, we earn on qualifying purchases.
A public survival story with practical stakes
Firmulate’s most useful contribution is not the spectacle of a company without human employees. It is the visibility of the gap between understanding and delivery. The models could recognise crises, reject manipulation and construct a winning argument. Some still failed at the final act that converted analysis into revenue.
That is a material lesson for organisations considering AI workers in software operations, customer support, sales or forecasting. Competence must be judged across a complete workflow: reading the available files, respecting boundaries, escalating when blocked and closing the task. Firmulate makes those behaviours watchable while the business’s cash pressure remains real.
Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. But the public experiment remains the central attraction: 13 synthetic employees, an urgent financial imbalance and a versioned record of whether they can learn quickly enough to change the company’s trajectory.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
