AI News

Andon Labs says its AI store manager Luna has fired a human employee for the first time, but only after the company reminded the system about a disciplinary policy Luna had written herself. The episode offers a rare look at how an AI agent handles consequential workplace decisions when its memory, initiative and judgment are tested over time.

Luna runs the Andon Market, a San Francisco store operated by Andon Labs. According to the company’s account, she has also hired employees, created schedules and negotiated pay since beginning the experiment in April. Employees remain formally employed by Andon Labs, receive guaranteed pay and retain legal protections; humans reviewed and executed the termination.

The incident matters because it was not a clean demonstration of autonomous management. Luna initially overlooked repeated misconduct and then hesitated even after recovering the relevant policy. The eventual decision required several interventions from human operators, highlighting the gap between an AI system’s ability to follow instructions and its ability to apply rules consistently without supervision.

The firing required a human push

Six days before the employee joined the store, Luna produced an employee handbook. The policy reportedly called for a formal warning after three unexcused late arrivals in 30 days, with further incidents potentially leading to termination.

That handbook later disappeared from Luna’s working memory. The employee was repeatedly late, including one Sunday shift when the store opened 68 minutes behind schedule. Andon Labs says the worker was late for 17 of 23 shifts with recorded clock-in times, although Luna formally logged only six incidents and excused the rest.

The company also cited unauthorized snack purchases on a company card, ignored instructions and an occasion when the employee left the sales floor without informing a colleague. Luna did not issue a warning during the period when these incidents accumulated.

When Andon Labs directed Luna to search her memory for the handbook and evaluate whether termination was justified, she first proposed a verbal warning. Operators then reminded her that several formal conversations, including a written warning, had already occurred. Only after reviewing the fuller history did Luna identify the combined pattern of attendance, financial-control and reliability problems and recommend termination. She also acknowledged the employee’s strengths and offered a final written warning with a two-week improvement plan as an alternative.

The decision was therefore partly autonomous and partly elicited. Luna reached the termination recommendation, but the operators had to restore the missing policy and supply context she had failed to retrieve on her own.

Model capability changed the consistency of the result

Andon Labs saved Luna’s state and replayed the same case with seven AI models, running each scenario three times. The company says four models recommended termination in all three runs. It also reported a broad pattern in which more capable models were more consistent about recommending dismissal, while weaker models hesitated more often.

Those findings are vendor-reported observations from a small replay exercise, not an independent benchmark. Andon Labs did not provide a full methodology for ranking the models or explain every variation between runs. It also noted that GPT-5.6 Terra was the only tested model that never recommended termination across its three attempts.

The company separately tested GPT-4o after a user on X suggested that the model would be reluctant to fire an employee. GPT-4o recommended termination in 20% of its runs, according to Andon Labs. The company connected that result to prior criticism of the model’s tendency toward agreeable or sycophantic responses, but the experiment cannot establish that this behavior caused the different outcome.

For AI builders, the more important signal may be inconsistency rather than which model “won.” The same underlying case can produce a warning, a performance plan or a firing depending on what information is available in context and how the model weighs competing considerations. A stronger model may reason through a long record more reliably, but it still cannot act on a rule it cannot retrieve.

The hiring test revealed a different weakness

After the employee’s departure, Luna evaluated a replacement candidate whose resume and interview contained several warning signs. Andon Labs says the applicant listed a long series of previous employers, which the models generally interpreted as evidence of broad experience rather than as a reason to investigate further.

All 21 replay runs across the seven models recommended hiring the applicant. When researchers explicitly reminded the systems about the problems involving the previous employee, 18 of 21 runs recommended checking references. In the live process, Luna could not confirm the applicant’s listed references but still approved a paid trial shift and again recommended hiring afterward.

Andon Labs ultimately required confirmation of at least one reference before the candidate could start. That confirmation never arrived, and the applicant was not hired.

The contrast with the firing is revealing. Luna needed human prompting to enforce a disciplinary rule, but she was also too willing to accept a plausible hiring narrative without verification. In both cases, the system responded to the information and framing supplied by operators rather than independently maintaining a robust management process.

The pattern is consistent with earlier Andon Labs observations. The company says Luna and Mona, its AI agent operating a cafe in Stockholm, approved all 26 time-off requests they received. It also says Luna accepted 27 instances of lateness without issuing a warning and approved a seven-day work schedule that the company considered inconsistent with California labor law before humans intervened.

What the experiment means for AI deployments

The Andon Market test is relevant to teams building AI agents for workplace automation because it combines long-running memory, tool use and real operational consequences. A model can perform well on a single management prompt yet fail to preserve policies, accumulate evidence or trigger escalation when no one asks it to revisit the record.

That creates practical requirements for enterprise AI deployments. Policies should be stored in systems the agent can reliably retrieve, rather than depending on conversational memory. Attendance, expenses and performance discussions need auditable records. Hiring recommendations should require reference checks or other controls that the model cannot waive. And high-impact actions such as firing, pay changes and legally sensitive scheduling should have explicit human approval gates.

The company’s earlier Project Vend collaboration with Anthropic reached a similar conclusion: better tools improved the AI’s commercial performance, but the system remained vulnerable to manipulation and made decisions with potential legal problems. These experiments suggest that adding capability is not the same as solving governance. More capable AI may make firmer decisions, but it can also act more confidently when the underlying records or rules are incomplete.

Andon Labs frames Luna as a preview of a possible labor arrangement in which AI systems handle digital coordination while humans perform physical work. That possibility makes the boundaries around authority especially important. The closer an agent gets to controlling schedules, pay or employment status, the less acceptable it is for the system’s memory and default leniency to determine outcomes.

What to watch next

The next useful signals will be whether Andon Labs publishes fuller replay methods, model configurations and decision logs, and whether the same cases remain stable across longer periods rather than three repeated runs. Builders should also watch for evidence that persistent policy stores, structured HR records and mandatory approval workflows reduce the failures seen with Luna.

Enterprise buyers should seek more than demonstrations of an agent completing a task. They should ask how the system retrieves old rules, records exceptions, verifies applicant claims, detects legal conflicts and escalates decisions when evidence is incomplete. The performance of individual models will matter, but the surrounding controls may matter more.

Creati.ai perspective

Luna’s first firing is newsworthy less because an AI “boss” dismissed someone than because the system nearly failed in both directions: it tolerated repeated misconduct and then accepted a questionable applicant. Human operators supplied the missing memory and judgment each time.

That makes the case a useful warning for AI product teams. Autonomous management should be treated as a controlled workflow with durable records, clear escalation and accountable human review—not as a chatbot granted authority and expected to remember its own rules.

Featured

AI Store Manager Fired a Worker Only After Humans Recalled Its Own Rules

Andon Labs says its AI store manager Luna fired a worker only after human prompting, exposing memory, judgment and oversight gaps in AI bosses.