This product was not featured by Product Hunt yet.
It will not be visible on their landing page and won't be ranked (cannot win product of the day regardless of upvotes).

Product upvotes vs the next 3

Waiting for data. Loading

Product comments vs the next 3

Waiting for data. Loading

Product upvote speed vs the next 3

Waiting for data. Loading

Product upvotes and comments

Waiting for data. Loading

Product vs the next 3

Loading

VeloxQuant-MLX

Run bigger local LLMs in less memory, on Mac

VeloxQuant lets you run bigger local AI models in less memory, fully private with no cloud required. One simple API compresses memory usage up to 16x while keeping generation fast on Apple Silicon.

Top comment

Hey Product Hunt 👋 I built VeloxQuant because I kept hitting the same wall running local models on my own Mac: the KV cache — not the model weights — is what actually kills your context length and your memory budget as generation goes on. Every "run an LLM locally" tutorial stops at loading a quantized model and calling `generate()`; almost none of them touch the thing that grows unbounded with every token you generate. So I built the piece that was missing: a compression engine purpose-built for the KV cache, with 43 methods pulled from the research literature (quantization, token eviction, cross-layer merging) behind one interface, backed by hand-written Metal kernels so the compression itself doesn't become the new bottleneck. A few things I'd love feedback on: - If you're already running local models (Ollama, LM Studio, raw MLX/llama.cpp), what's actually stopping you from going to longer contexts today — memory, speed, or tooling? - Anyone building offline-first or on-device products (healthcare, legal, fintech, defense) — what would make you trust a local inference stack enough to ship it, versus falling back to a cloud API? - I'm building this out into SDKs beyond Python (TypeScript, Go, Rust) plus a native macOS app and Android support — which of those would unlock something for you first? This is solo-built and open source (MIT), and it's an early step toward something I care about a lot: making private, on-device AI a default technical property of the hardware people already own, not something you have to opt into by giving up convenience. Happy to answer anything about the internals, the Metal kernels, or the roadmap.

About VeloxQuant-MLX on Product Hunt

Run bigger local LLMs in less memory, on Mac

VeloxQuant-MLX was submitted on Product Hunt and earned 2 upvotes and 1 comments, placing #155 on the daily leaderboard. VeloxQuant lets you run bigger local AI models in less memory, fully private with no cloud required. One simple API compresses memory usage up to 16x while keeping generation fast on Apple Silicon.

On the analytics side, VeloxQuant-MLX competes within Open Source, Developer Tools, Artificial Intelligence and GitHub — topics that collectively have 1.1M followers on Product Hunt. The dashboard above tracks how VeloxQuant-MLX performed against the three products that launched closest to it on the same day.

Who hunted VeloxQuant-MLX?

VeloxQuant-MLX was hunted by Rajveer Rathod. A “hunter” on Product Hunt is the community member who submits a product to the platform — uploading the images, the link, and tagging the makers behind it. Hunters typically write the first comment explaining why a product is worth attention, and their followers are notified the moment they post. Around 79% of featured launches on Product Hunt are self-hunted by their makers, but a well-known hunter still acts as a signal of quality to the rest of the community. See the full all-time top hunters leaderboard to discover who is shaping the Product Hunt ecosystem.

For a complete overview of VeloxQuant-MLX including community comment highlights and product details, visit the product overview.