Web search engine built from the CommonCrawl corpus
In the maker’s words
Hello HN, I'm building an alternative web search engine. You can try the demo here: https://search.xorsoft.dev/ My goal was to reach 1% of Google's index. This is 4 billion pages across 33.5 million domains. I was inspired by the earlier attempts shared here: "Building a web search engine from scratch with 3B neural embeddings" [1] and "Crawling a billion web pages in just over 24 hours" [2]. I decided not to replicate them and instead used CommonCrawl corpus, built a pipeline, and threw the results into a classic full-text search index. First, I got my hands on a dedicated Hetzner server with 4x10TB HDDs. It allowed me to crunch CommonCrawl data and stay within my hobby budget. The backend is served from a similar "home-grade" server but with NVMe disks instead. The compressed index weighs just over 2TB because I cap content length at 4KB per page (p50=2.9KB, p99=46KB) to fit it on disk…
Where people found it
- Hacker NewsShow HN: Web search engine built from the CommonCrawl corpus4 points3 comments2 days ago
More sites like Web search engine built from the CommonCrawl corpus
- Giving Opus 5.5 a simulated paint canvasstillwet.art
- AIHOT一个自己找热点、自己写日报的网站框架。把信源和精选标准换成你的,它就是你的行业热点站。
- Offrunmanage every coding agent from one workspace
- OpenDotsYour always-on AI coworkers that move between text, calls, and Slack.
- Pi podRun your pi coding agent in sandboxes on your own server
- Ledge.shRunnable Markdown Notes