Hi HN. We built a QA platform for AI agents.My cofounder spent 6 years at Veeva Systems building QA frameworks for regulated software. I spent 5 years in sales helping AI-native startups with their growth. We've both seen the issues companies face when deploying agents without a structured way of testing and monitoring them. The engineering teams we interviewed prior to building said they are doing manual spot-checks or just using LLM as a judge to check discrepancies.We generate synthetic test environments for 35 platforms (Salesforce, Jira, Stripe, Zendesk, Datadog, and 29 more). Each environment has ~200 adversarial queries with computed ground truth across 7 categories: clean lookups, ambiguous questions, multi-step operations, scope boundary tests, contradictory inputs, invalid assumptions, and context-dependent questions.The ground truth is initially computed by running SQL against the synthetic dataset so it's not just guessed by an LLM.We also do adversarial testing for pre-pro...
Want to discover more AI signals like this?
Explore Steek