splash
A local inference engine for Apple silicon, built around the model.

About splash
Splash is our open-source inference engine for Apple silicon, built around the model rather than around a model zoo. On a 48 GB M5 Pro it generates 210 tokens/s on Qwen3.6-35B-A3B, reopens a cached 32K context in 123 ms, and serves four concurrent requests at 357 tokens/s combined.
Where people found it
- Hacker NewsSplash: A Local Engine Built Around the Model5 points1 comments15 days ago
- GitHubsplash: A local inference engine for Apple silicon, built around the model.1.2k stars15 days ago
- Hacker NewsDFlash 2: Keep Drafting Parallel99 points18 comments2026-08-19
- Hacker NewsDFlash 2: Keep Drafting Parallel2 points0 comments2026-08-19
- Hacker NewsDFlash 2: Keep Drafting Parallel15 points1 comments2026-08-19
More sites like splash
- Giving Opus 5.5 a simulated paint canvasstillwet.art
- AIHOT一个自己找热点、自己写日报的网站框架。把信源和精选标准换成你的,它就是你的行业热点站。
- Offrunmanage every coding agent from one workspace
- OpenDotsYour always-on AI coworkers that move between text, calls, and Slack.
- Pi podRun your pi coding agent in sandboxes on your own server
- Ledge.shRunnable Markdown Notes