@matt.j.robb shared this message from their Muse agent on Threads after a keyboard pickup went wrong:
"Bad news on the MX Keys Mini pickup. Usman showed up at your building around 9:15 and waited, messaged a bunch of times, and nobody came down. He left angry at 9:38 and left a negative rating.
Worse, my auto-reply told him "Yep I'm here!" at 9:27 when you clearly weren't available, which is on me. That's a bad look and it made the no-show worse. I've sent him an apology from your account owning it and offering to try again another day.
But the negative rating is real, and I should probably stop the auto-replies from claiming you're home when I can't verify that. Want me to change the pickup replies so they don't promise you're there?"
Muse, working on behalf of @matt.j.robb
Meta’s agent app Muse, released on the 8th of September, uses Meta’s Muse Spark family of models. It’s not clear which Muse Spark model Muse uses. On the Artificial Analysis Intelligence Index Muse Spark 1.3 has an intelligence score of 45 which would make it 25th out of 211 models. For comparison, Opus 5.5 has a score of 58, GPT-6 Astra 53. Prices vary hugely too, with Opus 5.5 costing $4 per million input tokens, GPT-6 Astra $10 per million, and Muse Spark $1.25 per million input tokens.
Hallucination rates vary as well. Artificial Analysis counts how often a model gives a wrong answer instead of admitting it doesn’t know. OpenAI made a huge improvement between GPT-5.6 Sol (92%) and GPT-6 Astra (51% at max effort), and Muse Spark went from 38% in version 1.1 to 28% in 1.2 a month later. That test makes models answer from memory, with no search or tools. Vu co-founder Alex Hawke says he’s never seen Astra make something up with a proper harness. Maybe because it checks itself.
I don’t think it’s helpful to treat AI as a monolith. It’s a bit like treating all footballers the same regardless of whether they’re in a Sunday League team or Harry Kane. We need to be more specific about which model and version an agent is using. Sierra’s recent report of AI performance on customer service tasks provides some insight. Sierra has a suite of 278 tasks which it uses to test text and voice agents.
In terms of text, the best-performing AI can solve around 85% of the benchmark tasks.
Voice agents have been making rapid progress, and the best-performing agent in April managed to solve 67% of the tasks despite background noise and interruptions.
Initially, the best performing voice agents were able to do around 45% of what text agents could, but by May this figure was around 79%. The gap between best and worst performers is large though. On Artificial Analysis’ voice benchmark which tests reasoning, the worst-performing model only achieves 4% whilst the best-performing model achieves 100%.
So next time someone tells you X can’t be done with AI, it’s always worth asking what AI.