This product was not featured by Product Hunt yet. It will not be visible on their landing page and won't be ranked (cannot win product of the day regardless of upvotes).
Product upvotes vs the next 3
Waiting for data. Loading
Product comments vs the next 3
Waiting for data. Loading
Product upvote speed vs the next 3
Waiting for data. Loading
Product upvotes and comments
Waiting for data. Loading
Product vs the next 3
Loading
Inferbench
Benchmarks local LLM engines on your hardware
InferBench is a vendor-neutral CLI tool that benchmarks local LLM inference engines like omlx and llama.cpp directly on your own hardware. Instead of relying on external metrics, it reports real, measured tokens per second. InferBench auto-detects engines, runs a fixed prompt set, and recommends the fastest configuration for your setup. It is self-hosted and Apache 2.0 licensed.
The inspiration for InferBench stemmed from the constant need for absolute certainty regarding local LLM performance. Developers often have to rely on benchmark numbers generated on completely different hardware configurations, which rarely reflect actual, real-world performance. The primary problem to solve was this lack of reliable, local benchmarking. While existing solutions offer theoretical memory fit estimates or isolated, single-engine metrics, there was a distinct need for a tool that reports real, measured tokens per second directly on a user's own machine.
To address this, the approach evolved into creating a solution that is entirely vendor-neutral and cross-engine. Instead of just running simple tests, InferBench was engineered to automatically detect installed engines (like omlx and llama.cpp) and process a fixed prompt set through one shared HTTP harness. This evolution ensured that the tool not only benchmarks but actively recommends the absolute fastest configuration for any specific hardware setup.
Inferbench was submitted on Product Hunt and earned 2 upvotes and 1 comments, placing #38 on the daily leaderboard. InferBench is a vendor-neutral CLI tool that benchmarks local LLM inference engines like omlx and llama.cpp directly on your own hardware. Instead of relying on external metrics, it reports real, measured tokens per second. InferBench auto-detects engines, runs a fixed prompt set, and recommends the fastest configuration for your setup. It is self-hosted and Apache 2.0 licensed.
On the analytics side, Inferbench competes within Open Source, Developer Tools, Artificial Intelligence and GitHub — topics that collectively have 1.1M followers on Product Hunt. The dashboard above tracks how Inferbench performed against the three products that launched closest to it on the same day.
Who hunted Inferbench?
Inferbench was hunted by Sourav Nandy. A “hunter” on Product Hunt is the community member who submits a product to the platform — uploading the images, the link, and tagging the makers behind it. Hunters typically write the first comment explaining why a product is worth attention, and their followers are notified the moment they post. Around 79% of featured launches on Product Hunt are self-hunted by their makers, but a well-known hunter still acts as a signal of quality to the rest of the community. See the full all-time top hunters leaderboard to discover who is shaping the Product Hunt ecosystem.
For a complete overview of Inferbench including community comment highlights and product details, visit the product overview.
The inspiration for InferBench stemmed from the constant need for absolute certainty regarding local LLM performance. Developers often have to rely on benchmark numbers generated on completely different hardware configurations, which rarely reflect actual, real-world performance. The primary problem to solve was this lack of reliable, local benchmarking. While existing solutions offer theoretical memory fit estimates or isolated, single-engine metrics, there was a distinct need for a tool that reports real, measured tokens per second directly on a user's own machine.
To address this, the approach evolved into creating a solution that is entirely vendor-neutral and cross-engine. Instead of just running simple tests, InferBench was engineered to automatically detect installed engines (like omlx and llama.cpp) and process a fixed prompt set through one shared HTTP harness. This evolution ensured that the tool not only benchmarks but actively recommends the absolute fastest configuration for any specific hardware setup.
Repo: https://github.com/RudrenduPaul/InferBench
MCP Servers:
https://mcpservers.org/servers/rudrendupaul/inferbench
https://glama.ai/mcp/servers/RudrenduPaul/inferbench
NPM: https://www.npmjs.com/package/inferbench-cli
PyPI: https://pypi.org/project/inferbench-cli