The news
David Pogue ran the new AI version of Siri through 125 tests after using the beta since shortly after WWDC. The final release is now available, and his results show a handful of outright failures alongside several tasks that performed beyond expectations. Pogue concludes that the biggest barrier is not the model’s accuracy but the user habit of opening apps instead of speaking to Siri.
Context
Siri has long required users to navigate menus and launch specific applications for routine actions. The new model in iOS 27 is intended to let people describe goals in natural language and have the system carry them out across apps. Pogue spent the summer on the beta and grew increasingly confident in the approach before running the structured test set on the shipping version.
Details
Pogue selected tasks from Apple’s own claims, ideas posted on Reddit, and scenarios he wondered about himself. The tests covered both straightforward commands and more open-ended requests. A small number of attempts failed to produce the expected result. In other cases the assistant executed steps that Pogue had not anticipated it could handle.
He notes that success depends on users remembering the assistant exists rather than defaulting to manual interaction. “The hardest part will be getting out of the habit of opening apps, and remembering that you have Siri,” Pogue wrote. Once that adjustment occurs, he expects users to save time on repetitive actions that previously required multiple taps and context switches.
Pogue’s write-up does not list every test outcome, but he emphasizes that the model is capable of actions most people would not assume fall within Siri’s scope. The evaluation was performed on the release version of iOS 27, not the earlier beta builds.
Why it matters
For people who already carry an iPhone, the practical change is modest in scope but potentially large in daily friction. A system that can reliably chain actions across applications removes the need to remember exact menu paths or switch contexts repeatedly. The remaining failures indicate that the model still requires verification on complex or ambiguous requests, so users cannot treat every response as final.
The larger shift is behavioral. Apple’s previous Siri encouraged short, narrow commands because that was the limit of its capability. The new version invites longer, goal-oriented speech, which only works if people stop reaching for the home screen first. Pogue’s experience suggests the time savings appear once that habit forms, but the transition itself may take weeks of deliberate effort.
Whether the feature reaches most users will depend on how consistently the model performs outside the 125 tested cases. Early adopters willing to experiment will discover both the ceiling and the remaining gaps. Everyone else will continue opening apps until the assistant proves itself in daily use.
---
Sources:
No comments yet