Run eval experiments at scale in realistic environments. Define custom task sets to build your private benchmarks, measure how well agents can use any product, and find best models for your use cases. Generate dynamic insights to detect frictions in product interfaces or token inefficiencies.
Hey Product Hunters, I’m Haritha, co-founder of Oqoqo
Every week there is a new model launch and yet another benchmark released in the wild. But they do not help product builders evaluate how well their products can be discovered and used by these agents and models, or talk about actual tasks their users would perform. Most benchmarks today exist in curated environments and do not translate well to the real world.
We built Oqoqo to bridge this gap. Oqoqo makes it super simple to build realistic evals and custom benchmarks for tasks users actually care about.
With Oqoqo, you can define tasks as simple as a prompt your user might give to an agent e.g. “integrate supabase to my webapp to store user sign ups”, provide what you want to test for e.g. Supabase SDK, API, CLI etc. and define what success looks like e.g. “must set up RLS”.
We handle the rest. Our infrastructure spins up isolated sandboxes, executes the tasks against agents of your choice, catalogs every single step the agents take including tool calls, retries, discovery loops etc, and documents token consumption, cost, along with evaluating success/failure based on your success criteria.
With Oqoqo you can:
Reliably measure how agent friendly your product surfaces are against Codex, Claude Code, OpenClaw, Hermes, Pi, Opencode, Cursor, GitHub Copilot
Regression test MCP, CLI, skills, SDK, and any agent facing interface (we are continuously using Oqoqo to dogfood and improve our own MCP/CLI)
Create and share custom benchmarks for how agents discover and use your product
Compare models and harnesses for domain specific tasks
See whether new versions improve agent experience
We built Oqoqo for teams building products that agents want to use, and for teams putting agents into day to day work.
And the best thing? Your agent can handle the setup for you ✨, try it out for free today: https://oqoqo.ai/
We would love to learn what kind of experiments you would like to run and what questions you have about agent interactions and agent experience.
Congrats on the launch. Oqoqo looks like a really interesting approach to evaluating AI agents in realistic environments.
The idea of testing agents on the real world tasks make sense. How long does it take to setup an eval with oqoqo ?
the realistic environments part is the right fight. the thing id watch next is eval rot, a case written six months ago measures the world as it was the day someone wrote it, and a suite that stops failing looks exactly the same as a product that got good. the number id surface is what share of cases have ever failed, because the ones that never have arent tests, theyre decoration, and they pile up until the green means nothing
This is super useful! I’ve been building agents and skills to make product onboarding easier for enterprise customers but right now its kind of a black box - I don’t really know how they are using it, where things are breaking and what I should fix first. If I get to see how the agent behaves across diff scenarios and where users are getting stuck it will be huge. Can't wait to use the CLI and run this on autopilot!
Evaluating models and harnesses in an easy, consistent way is hard. Good to see your platform take up the challenge and ease the entire process. 10/10 recommend
the same task rarely takes the same path twice with an agent, different tool call order, different retries. how are you keeping the scoring stable run over run so a benchmark result doesn't just become noise from agent nondeterminism
The part eval systems often miss is recovery behavior: permission denial, stale credentials, partial side effects, and a rerun after failure. A benchmark that scores the happy path but not cleanup and recovery can reward an agent that looks finished while leaving the product in a worse state.
The token efficiency insights caught my attention. Small inefficiencies can become pretty expensive when agents run at scale.
Evaluation becomes a major challenge once AI systems move beyond demos. What experience pushed you toward building a dedicated platform for this problem?
How do you handle tasks where an agent technically completes the job but the quality of tge result is still poor?
Have you noticed big differences between Claude Code, Codex, Cursor, and Copilot when running the exact same real world task?
“Build evals and custom benchmarks for real-world tasks”
oqoqo launched on Product Hunt on August 10th, 2026 and earned 311 upvotes and 29 comments, earning #1 Product of the Day. Run eval experiments at scale in realistic environments. Define custom task sets to build your private benchmarks, measure how well agents can use any product, and find best models for your use cases. Generate dynamic insights to detect frictions in product interfaces or token inefficiencies.
oqoqo was featured in Software Engineering (42.8k followers), Developer Tools (517.3k followers) and Artificial Intelligence (475.7k followers) on Product Hunt. Together, these topics include over 199.2k products, making this a competitive space to launch in.
Who hunted oqoqo?
oqoqo was hunted by fmerian. A “hunter” on Product Hunt is the community member who submits a product to the platform — uploading the images, the link, and tagging the makers behind it. Hunters typically write the first comment explaining why a product is worth attention, and their followers are notified the moment they post. Around 79% of featured launches on Product Hunt are self-hunted by their makers, but a well-known hunter still acts as a signal of quality to the rest of the community. See the full all-time top hunters leaderboard to discover who is shaping the Product Hunt ecosystem.
Want to see how oqoqo stacked up against nearby launches in real time? Check out the live launch dashboard for upvote speed charts, proximity comparisons, and more analytics.
Hey Product Hunters, I’m Haritha, co-founder of Oqoqo
Every week there is a new model launch and yet another benchmark released in the wild. But they do not help product builders evaluate how well their products can be discovered and used by these agents and models, or talk about actual tasks their users would perform. Most benchmarks today exist in curated environments and do not translate well to the real world.
We built Oqoqo to bridge this gap. Oqoqo makes it super simple to build realistic evals and custom benchmarks for tasks users actually care about.
With Oqoqo, you can define tasks as simple as a prompt your user might give to an agent e.g. “integrate supabase to my webapp to store user sign ups”, provide what you want to test for e.g. Supabase SDK, API, CLI etc. and define what success looks like e.g. “must set up RLS”.
We handle the rest. Our infrastructure spins up isolated sandboxes, executes the tasks against agents of your choice, catalogs every single step the agents take including tool calls, retries, discovery loops etc, and documents token consumption, cost, along with evaluating success/failure based on your success criteria.
With Oqoqo you can:
Reliably measure how agent friendly your product surfaces are against Codex, Claude Code, OpenClaw, Hermes, Pi, Opencode, Cursor, GitHub Copilot
Regression test MCP, CLI, skills, SDK, and any agent facing interface (we are continuously using Oqoqo to dogfood and improve our own MCP/CLI)
Create and share custom benchmarks for how agents discover and use your product
Compare models and harnesses for domain specific tasks
See whether new versions improve agent experience
We built Oqoqo for teams building products that agents want to use, and for teams putting agents into day to day work.
And the best thing? Your agent can handle the setup for you ✨, try it out for free today: https://oqoqo.ai/
We would love to learn what kind of experiments you would like to run and what questions you have about agent interactions and agent experience.