Cekura is the testing, observability, and self-improvement platform for production voice and chat AI agents. It simulates thousands of scenarios, catches failures, diagnoses the root cause, rewrites prompts and config, then re-validates with a full regression sweep. Unlike tools that hand failures back to your team, Cekura closes the loop by fixing the agent itself and proving the fix holds without overfitting.
Today we are launching self-improving loops for voice agents.
Fixing a voice agent has always been fragmented. Your testing tool tells you what failed, you diagnose it from transcripts, patch the prompt, re-run, and something else breaks. The tools find problems, but the fixing has always been a human walking between them.
Cekura collapses that loop. It runs thousands of simulated calls and groups every failure, explained in plain English. Optimise agent hands them to the Cekura Agent, or any coding agent you use, Claude Code, Codex, anything. It reproduces each failure, makes the change, and reruns until everything passes, then verifies nothing else broke. You review the diff.
Two rules: it must reproduce a bug before fixing it, and every fix is proven on simulated calls on a clone, never your live agent.
Free for everyone to try, starting today. We are in the comments all day. If you want to chat more, please feel free to book time here
Voice AI quality is surprisingly difficult to measure consistently. I like that this goes beyond transcript evaluation and looks at things like interruptions, latency, and delivery quality. That's usually where real production issues show up
Love seeing more attention on production reliability for voice AI. The regression validation especially stood out. Curious whether Cekura can prioritize issues by business impact (for example, failed payments vs. minor conversation hiccups) or if everything is treated equally.
Congrats on the launch! Really like that you’re focusing on proving fixes instead of just identifying failures. One thing I’m curious about: how do you decide when a suggested fix is reliable enough to recommend versus flagging it for manual review?
Love that this closes the loop instead of dumping another list of flagged calls on a human. The rule that it has to reproduce a bug before it is allowed to fix it feels very sane. For a team running a voice agent in something sensitive like healthcare intake, can you control which kinds of fixes it applies on its own and which ones need sign-off first?
Voice agents fail in ways text evals never catch — interruptions, latency, someone talking over the bot. The "loop" framing is the right one: testing voice once at build time is basically useless. Would love to know how many simulated calls it takes before the improvements show up.
Thanks — the GitHub Actions path is the answer to that question. The follow-up I'd have is whether the infra regression suite runs against a live agent clone or replays recorded sessions, because customer-support agents with integration state tend to behave differently on replay versus a live environment.
Love seeing the focus on proving that a fix doesn't introduce regressions elsewhere. Reliable AI systems need repeatable validation, not just faster debugging.
Simulating messy real-world conversations with interruptions, pauses, and background noise feels much closer to production than traditional scripted evaluations.
Does Cekura provide regression testing capabilities to catch quality drops when agents are updated or retrained?
How customizable are the evaluation metrics? Can teams define their own quality benchmarks based on their specific use case?
For teams building voice based agents specifically, does Cekura account for latency and audio quality issues, or is the focus primarily on conversational logic?
How does Cekura evaluate more nuanced qualities like tone, empathy, or conversational flow, beyond just accuracy or task completion?
Does Cekura support testing across multiple languages and accents, given how critical that is for conversational AI reliability?
Congrats on the launch. The hard part of a self-improvement loop is the blast radius. When an agent rewrites its own behavior from production calls, a fix for one flow can quietly bend a compliance flow sitting next to it, and in healthcare that is the thing every security review hunts for. Closing that loop safely, so the agent improves without drifting out of its guardrails, is the whole game, and it is a genuinely hard problem to have taken on.
When monitoring production calls, how does Cekura detect quality issues in real time, and what kind of alerting or reporting does it provide to teams?
About Cekura on Product Hunt
“The self-improvement loop for voice agents”
Cekura launched on Product Hunt on July 28th, 2026 and earned 336 upvotes and 68 comments, earning #2 Product of the Day. Cekura is the testing, observability, and self-improvement platform for production voice and chat AI agents. It simulates thousands of scenarios, catches failures, diagnoses the root cause, rewrites prompts and config, then re-validates with a full regression sweep. Unlike tools that hand failures back to your team, Cekura closes the loop by fixing the agent itself and proving the fix holds without overfitting.
Cekura was featured in SaaS (43.3k followers), Developer Tools (516.5k followers) and Audio (2.1k followers) on Product Hunt. Together, these topics include over 132.4k products, making this a competitive space to launch in.
Who hunted Cekura?
Cekura was hunted by Garry Tan. A “hunter” on Product Hunt is the community member who submits a product to the platform — uploading the images, the link, and tagging the makers behind it. Hunters typically write the first comment explaining why a product is worth attention, and their followers are notified the moment they post. Around 79% of featured launches on Product Hunt are self-hunted by their makers, but a well-known hunter still acts as a signal of quality to the rest of the community. See the full all-time top hunters leaderboard to discover who is shaping the Product Hunt ecosystem.
Want to see how Cekura stacked up against nearby launches in real time? Check out the live launch dashboard for upvote speed charts, proximity comparisons, and more analytics.
Hey Product Hunt!
Sidhant here, co-founder of Cekura
Today we are launching self-improving loops for voice agents.
Fixing a voice agent has always been fragmented. Your testing tool tells you what failed, you diagnose it from transcripts, patch the prompt, re-run, and something else breaks. The tools find problems, but the fixing has always been a human walking between them.
Cekura collapses that loop. It runs thousands of simulated calls and groups every failure, explained in plain English. Optimise agent hands them to the Cekura Agent, or any coding agent you use, Claude Code, Codex, anything. It reproduces each failure, makes the change, and reruns until everything passes, then verifies nothing else broke. You review the diff.
Two rules: it must reproduce a bug before fixing it, and every fix is proven on simulated calls on a clone, never your live agent.
Free for everyone to try, starting today. We are in the comments all day. If you want to chat more, please feel free to book time here