Cactus Needle 3
8-29MB automation models can match DeepSeek V4 Flash
About Cactus Needle 3
A foundation model for mobile, wearables, robots, smart home, automotive and microcontrollers. One 8-29 MB binary that beats models 10x its size on mobile tool calls and matches 2-3x bigger models on extraction.
In the maker’s words
Hey HN, Henry from Cactus here. We submitted Needle 2 here a few weeks ago, and the feedback in the discussion thread was incredibly valuable, thanks! Thanks to all that feedback, we’ve been able to move quickly to release Needle 3 and I'd love to hear what you think again. The key features: 1) Automation (tool calls & structured JSON output): Needle still doesn't chat by design, its quite challenging to pack general capacity into such small models, so we focus on tool calls and structured JSON. If no tool you declared fits the request, you get an empty list back (note for when playing with the demo). 2) Intelligence Laddering: Every layer (2 to 20) is a deployable subnetwork, so one set of weights, 25 to 121 million parameters at 2-bit, shipping as 8-29MB binaries. On a Raspberry Pi 5 it decodes at up to 4k tokens/sec and prefills at up to 10k. 3) Monarch Hadamard MLP: replaces the dens…
Where people found it
- Hacker NewsWhistle: Speech to Text in 16.9 MB4 points1 comments1 day ago
- Hacker NewsWhistle: Speech to Text in 16.9 MB by Cactus Compute2 points0 comments1 day ago
- Hacker NewsShow HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash236 points93 comments16 days ago
- Hacker NewsShow HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots537 points185 comments2026-08-10