• News In Brief
  • Influence Excellence Awards 2026
  • AI
  • Education
  • Pro AV
  • Case Study
  • Interview
No Result
View All Result
SUBSCRIBE
Smart Solutions World
  • News In Brief
  • Influence Excellence Awards 2026
  • AI
  • Education
  • Pro AV
  • Case Study
  • Interview
No Result
View All Result
No Result
View All Result
Home AI

KushoAI Benchmark Finds AI Coding Tools Struggle With Complex API Bugs

SmartSolutionUser1 by SmartSolutionUser1
June 10, 2026
in AI
0
KushoAI Benchmark Finds AI Coding Tools Struggle With Complex API Bugs
76
SHARES
1.2k
VIEWS
Share on FacebookShare on Twitter

KushoAI released the first comparative benchmark study of how leading AI coding and testing agents perform at finding bugs in live APIs. While AI tools generate plausible tests quickly, most struggle to detect bugs emerging from field relationships, operation semantics, and business-logic dependencies.

You might also like

Defining the Future of AI Security – Akamai Selected as Strategic Security Partner for WWT’s ARMOR Framework

Introducing GPT-Live & a new ChatGPT Voice experience

The B2B Marketer’s AI Reckoning – Own the Disruption or Fall Behind  

The report evaluated seven AI systems across three groups: general-purpose LLMs, coding agents, and KushoAI’s API testing agent. Each received only a JSON schema and a sample payload for 20 live API scenarios, each containing 97 known functional bugs across three difficulty tiers.

The central finding is a sharp drop in performance as bugs get more complex. Most systems catch simple schema violations: missing fields, wrong types, and null values. Performance falls when detection requires semantic reasoning or understanding how valid fields combine into an invalid business state. On the hardest tier, the strongest coding-agent workflow detected 53%, the strongest general-purpose LLM detected 34%, and KushoAI detected 76%, ranking first across every complexity tier.

Mr. Abhishek Saikia, Co-founder and CEO of KushoAI
Mr. Abhishek Saikia, Co-founder and CEO of KushoAI

“AI can generate tests. That is no longer the hard question,” said Mr. Abhishek Saikia, Co-founder and CEO of KushoAI. “The harder question is whether those tests reach the failure modes that matter. Simple schema-level testing is increasingly table stakes. The real gap appears when API testing requires reasoning across fields, states, and business rules.”

This report follows KushoAI’s earlier launch of APIEval-20, the industry’s first open benchmark for evaluating AI agents on API bug detection from schema and payload alone. This study reveals how general-purpose LLMs, coding agents, and purpose-built API testing agents actually perform.

Better prompting helps but does not close the gap. Prompt chaining improved field-level coverage but did not produce the cross-field tests needed to catch business-logic failures. KushoAI showed the lowest run-to-run variance, critical for teams integrating generated tests into CI pipelines.

The findings build on KushoAI’s analysis of 1.4 million test executions across 2,616 organizations. The report positions APIEval-20 as an emerging standard, similar to the role HumanEval and SWE-bench play in software engineering research.

If you have an interesting Article / Report/case study to share, please get in touch with us at editors@roymediative.com roy@roymediative.com, 9811346846/9625243429

Tags: Benchmark Finds AICoding ToolsComplex API BugsKushoAIsmart solutions world
Share30Tweet19
SmartSolutionUser1

SmartSolutionUser1

Recommended For You

Defining the Future of AI Security – Akamai Selected as Strategic Security Partner for WWT’s ARMOR Framework

by SmartSolutionUser1
July 10, 2026
0
Defining the Future of AI Security – Akamai Selected as Strategic Security Partner for WWT’s ARMOR Framework

Akamai announced its selection as a strategic partner for World Wide Technology’s (WWT) AI Readiness Model for Operational Resilience (ARMOR). This collaboration positions Akamai as a foundational security...

Read moreDetails

Introducing GPT-Live & a new ChatGPT Voice experience

by SmartSolutionUser1
July 10, 2026
0
Introducing GPT-Live & a new ChatGPT Voice experience

OpenAI introduced GPT-Live, a new generation of models powering ChatGPT Voice. GPT-Live is our smartest voice model yet, and is built on research advances that make conversations with...

Read moreDetails

The B2B Marketer’s AI Reckoning – Own the Disruption or Fall Behind  

by SmartSolutionUser1
July 10, 2026
0
The B2B Marketer’s AI Reckoning – Own the Disruption or Fall Behind  

Marketing has never stood still, and every era has brought with it a force that challenges established playbooks. Today, that force is AI. For marketers, AI has sparked...

Read moreDetails

Nurix AI to Acquire Verloop.io, Expanding Conversational AI Platform Across Voice and Chat

by SmartSolutionUser1
July 10, 2026
0
Nurix AI to Acquire Verloop.io, Expanding Conversational AI Platform Across Voice and Chat

Nurix AI, an enterprise AI company building autonomous AI software and agents for complex business workflows, announced the acquisition of Verloop.io, one of India’s leading enterprise conversational AI...

Read moreDetails

MathCo Announces Second Edition of Spark AI, Bringing Together Enterprise AI Leaders and Practitioners

by SmartSolutionUser1
July 9, 2026
0
MathCo Announces Second Edition of Spark AI, Bringing Together Enterprise AI Leaders and Practitioners

MathCo, a global leader in Enterprise AI, announced the second edition of Spark AI 2026, its flagship knowledge-sharing platform, scheduled for July 17, 2026. Bringing together enterprise AI...

Read moreDetails
Next Post
Nagarro partners with BrowserStack to ​​supercharge AI-powered testing workflows for enterprises

Nagarro partners with BrowserStack to ​​supercharge AI-powered testing workflows for enterprises

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Browse by Category

Browse by Category

Smart Solutions World

We bring you the best Premium news, magazine, personal blog, etc. Check our landing page for details.

  • News In Brief
  • Influence Excellence Awards 2026
  • AI
  • Education
  • Pro AV
  • Case Study
  • Interview

BROWSE BY TAG

Agentic AI Agora AI Akamai AMD Cloudflare CloudKeeper Coforge CrowdStrike Cybersecurity Databricks Fortinet Gartner GenAI Google Cloud HCLTech Hitachi Vantara Honeywell IBM Infosys Kaspersky Keysight Kramer LTIMindtree Microsoft New Relic Nvidia OpenAI Palo Alto Networks PPDS Qlik Qualcomm Seqrite ServiceNow SiMa.ai smart solutions world smartsolutionsworld smart solutions world latest news Software Tata Communications Tech Mahindra Technology Tenable UiPath Vertiv

© 2024 NCN - Premium news & magazine by NCN.

No Result
View All Result
  • News In Brief
  • Influence Excellence Awards 2026
  • AI
  • Education
  • Pro AV
  • Case Study
  • Interview

© 2024 NCN - Premium news & magazine by NCN.

Not enough quota to unlock this post
Unlock left : 0
Are you sure want to cancel subscription?