AI News Feed
Market watch
Products & Applications

Decodo study finds AI agents can't fully master real-world browsing tasks

None of 45 AI agents aced Decodo real-world browsing tests; Claude for Chrome top-scored with 18/20.

Claude for Chrome was the highest-scoring agent, reaching 18 points out of 20. Its OpenAI counterpart, the ChatGPT Chrome Extension, scored 14 points. Decodo found that transactions were the weakest category, with an average score of 0.43 out of 2. While agents could often reach the checkout stage, they could not complete a purchase on the user's behalf. The research also warned that even when an agent could technically process a purchase, it might not have adequate safeguards to protect sensitive information such as credit card numbers.

Safety measures were a recurring theme in the study. More than half of the agents capable of multi-step workflows lacked documented safeguards before carrying out irreversible actions. However, browser-native agents in the test received full marks for cross-tab awareness. Because performance varied by category, Decodo advised users to select an agent based on their planned usage rather than relying on long feature lists. Gabriele Vitke, product marketing team lead at Decodo, said: “Match the tool to the job, not to the longest feature list,” urging AI agent users to focus on their own needs instead of being swayed by marketing.