Tag Archives: Parallel computing
texrex web page cleaning system
Moved to GitHub as of 1 May 2016 (from SourceForge rev. 622).
This is the work horse web page cleaning system behind the COW. It turns crawled HTML documents into clean XML corpus documents. It is released under a permissive 2-clause BSD license. Continue reading
Processing and Querying Large Web Corpora with the COW14 Architecture (2015)
Roland Schäfer. Processing and Querying Large Web Corpora with the COW14 Architecture. In Proceedings of Challenges in the Management of Large Corpora (CMLC-3) (IDS publication server). 28–34. [BibTeX]