Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Hi. I work for commoncrawl. We are about to start an improved recrawl and will be doing this more frequently going forward. In the process we will also consolidate our data on S3 to keep it relevant. But, as with any crawl of the Internet, there is lot of noise in there. We spent most of 2011 tweaking the algorithms to improve the freshness and quality of the crawl, and hopefully this work starts to show results in 2012.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: