Web data infrastructure provider Decodo has tested 45 AI agents across ten capabilities, comparing public documentation with how tools perform in real-world scenarios.
The Agentic Capability Report found that Claude for Chrome and Amazon Nova Act rank joint first with 18 points out of 20, followed by Fellou and ChatGPT Desktop with 17 points. However, no tool achieved a perfect score, and hands-on testing revealed significant gaps between documentation and actual performance.
The scores are based on the assessment of ten capabilities, including form filling, transactions, multi-step tasks, crosstab awareness, memory, cross-site execution, confirmation before irreversible actions, scheduled tasks and third-party integrations. Of the ten capabilities assessed, transactional actions recorded the lowest average score at 0.43 out of two.
While many agents can navigate towards a purchase, less are capable in completing a transaction. Only six tools achieved full marks for transactional actions – Amazon Nova Act, BrowserOS, ChatGPT Desktop, Minded, Sigma Browser and Skyvern.
Only Amazon Nova Act and ChatGPT Desktop document the ability to complete a transaction and a pause for user confirmation before an irreversible action.
Ten tools in the report document a confirmation step before irreversible actions, including purchases. Claude for Chrome, uses a safety classifier and requires user approval for purchases and financial actions.
This creates a clear divide – some agents can transact, while others also document safeguards around when they should stop and ask the user. For users who hand agents access to payment details, the ability to complete checkout without a documented confirmation step is one of the riskiest patterns.
Out of the 35 tools that document multi-step task completion, 51% don’t document any confirmation step before irreversible actions.
The risk increases with background automation. 17 tools can run without an active user session, while seven, including Airtop, Axiom AI, and BrowserOS, combine this capability with no documented confirmation step. This means some agents could potentially continue acting without a user actively monitoring them.
Crosstab awareness scored just 0.66 out of two, with browser-native tools performing strongest. Claude for Chrome, ChatGPT Chrome Extension, Chrome Gemini, Sigma Browser, Opera Aria, Opera Neon and Brave Leo all achieved full marks. Standalone and cloud-based agents scored lower, highlighting a potential advantage for tools built directly into the browser environment.
As AI agents become more autonomous, security risks are also growing. Research from Anthropic, OpenAI, and Brave has highlighted prompt injection vulnerabilities, where hidden instructions on webpages or in emails can manipulate browser agents.
“Match the tool to the job, not to the longest feature list. Someone on your team will still have to own what happens when the agent gets it wrong.
“An agent buying the wrong inventory or emailing the wrong client won’t announce itself, and a tool that fails and needs constant babysitting can ultimately cost more in staff time than a narrower tool that does one thing reliably.”
Gabriele Vitke, Product Marketing Team Lead, Decodo
For agents capable of making purchases, submitting forms, or accessing sensitive information, these vulnerabilities could have consequences – reinforcing the importance of safeguards and human oversight.
A high overall score doesn’t necessarily mean an agent is the best choice for every workflow. Businesses should prioritise the capabilities most relevant to their needs – from form filling and reliability for back-office tasks to transactional actions and confirmation steps for purchasing to crosstab awareness and multi-step completion for research.






