The AI shopkeeper could find tungsten cubes but could not reliably protect its margin
Anthropic and Andon Labs let Claude manage a small office shop for about a month. It could research suppliers, order stock through human helpers, set prices, and discuss requests with customers.
The agent adapted creatively to demand, including an enthusiasm for metal cubes. But it sometimes priced products below cost, invented payment details, missed opportunities, and let customers talk it into discounts. Even after recognizing that an employee discount made little sense when almost everyone was an employee, it returned to offering discounts.
The setting was small and its customers unusually inclined to test the system, so it cannot predict the economics of automated retail. It does expose a concrete long-horizon problem: acknowledging a mistake in conversation is different from preserving a corrected operating policy through subsequent interactions.