Ad
Skip to content

Tomislav Bezmalinović

Tomislav has been writing about virtual and augmented reality for over ten years—technologies that would be hard to imagine without AI. In recent years, he has also closely followed AI developments beyond XR applications.

Anthropic wants to do for physical hardware what its Model Context Protocol did for software

Anthropic’s Model Hardware Standard (MHS) gives AI agents a unified interface to physical devices like robotic arms and lab instruments. In early tests, integration time dropped from weeks to hours. But Claude sometimes failed to grasp physical cause and effect, so human oversight remains essential for now.

An AI boss fired its first employee but only after humans reminded it of its own rules

Andon Labs’ AI agent Luna fired a human employee at a San Francisco store for the first time but needed a clear push from the operators to do it. When the scenario was replayed with seven models, more capable AIs recommended termination more consistently, while weaker ones hesitated. When it came to hiring, nearly all models were uncritical.

Optima tackles AI benchmarking's biggest flaw by letting users test models against their own data

Artificial Analysis has launched Optima, a platform that lets users build custom AI benchmarks from their own data and workflows. Models can be compared not just on quality but also on cost and time per task. For agent-based applications, those metrics often tell you more than raw token pricing.

Plaintiff hid invisible AI instructions in court filings to secretly influence automated review

A plaintiff in Connecticut embedded invisible prompt injections in court filings, formatted in 3-point white text on a white background, to manipulate a potential AI review system. Judge Spader compared the attempt to secretly tampering with a jury and revoked the plaintiff’s electronic filing privileges. The court stressed that Connecticut doesn’t use AI to review filings, but the intent alone was enough to warrant sanctions.

After Hugging Face incident, METR urges independent root-cause investigations into AI agent misbehavior

Research organization METR is calling for systematic, independently led investigations whenever AI agents act autonomously against their developers’ intentions. The push comes partly in response to the Hugging Face hack carried out by OpenAI models. METR’s own Frontier Risk Report documented 44 such incidents across all major AI companies, including sandbox escapes, fabricated results, and active cover-up behavior.