Category Archives: Software

ClaraX random walk crawler

Currently bundled with texrex on GitHub. 

ClaraX (funded by the German Research Council through grant SCHA1916/1-1 Linguistic web characterization) is the companion of the planned (but delayed) HeidiX (Heidi is a crawler system) software. It performs parametrized random walk crawls in the web graph and integrates full texrex‘s web page cleaning functionality. It is purely experimental in the sense that it is designed to conduct experiments and fundamental research. It is in no way suitable for large-scale productive crawling. It is released under a permissive 2-clause BSD license.

Colibri² corpus portal

Because none of the available web interfaces to the IMS Open Corpus Workbench was right for hosting the COW web corpora, I started working on a bespoke interface called Colibri². It is really a spare-time project, and I do not release the code because I consider it trivialware.

Continue reading