Repository for the GreenNLP/OpenEuroLLM document descriptor research project. The project uses large language models to create a dynamic taxonomy of descriptive labels ("descriptors") for web documents. Descriptors can be used to filter, subsample, and retrieve relevant documents from a much larger collection.
The code was developed for the LUMI supercomputer (LUMI) and is most useful as a research artifact and starting point for further experiments. It is not a turnkey, platform-independent package.
Interactive demo: http://159.69.54.25:5000/. The demo is external to this repository and may no longer be available.
Data release on Hugging Face: https://huggingface.co/datasets/TurkuNLP/WebDocumentDescriptors.
descriptor_generation/generates descriptors from documents.disambiguation/groups and disambiguates descriptors.merging/merges synonymous descriptors and resolves remaining duplicates.harmonize/aligns descriptors with an existing schema.faiss/builds descriptor indexes and runs retrieval and LLM-based evaluation.LLM_as_judge/contains the judging and label-inference pipelines.eval_descriptors/contains descriptor-quality and growth analyses.notebooks/contains exploratory analyses; notebooks may require paths and outputs that are not committed.scripts/contains data preparation and utility scripts.requirements.txtrecords one Python environment used by the project.
The basic workflow is:
- Generate descriptors with
descriptor_generation/generate_descriptors.py. - Extract descriptor groups with
disambiguation/extract_descriptor_groups.py. Use--num-splitsto divide large inputs for parallel processing. - Disambiguate the groups with
disambiguation/disambiguate_descriptors.py. - If jobs were run in parallel, concatenate their outputs with
disambiguation/concat_disambig_results.sh. - Repeat extraction and disambiguation as needed.
- Merge synonyms with
merging/merge_synonyms.py. - If duplicates remain, force-merge them with
merging/force_merge.py.
The run_*.sh files beside the Python modules contain LUMI/Slurm examples and should be read before submitting a job.
- Generate descriptors with
descriptor_generation/generate_descriptors.py. - Align them with the schema using
harmonize/harmonize_with_schema.py.
faiss/full_pipeline.py runs the FAISS search, judges descriptor matches, evaluates surviving documents, and writes JSONL and summary artifacts. The supplied Slurm wrapper is faiss/run_full_pipeline.sh; it currently expects format or topic and raw or harmonized arguments and contains site-specific paths that must be adapted.
For a new environment, inspect the argument parser in faiss/full_pipeline.py and the lower-level scripts in faiss/ before adapting the command. The previous README example referred to a non-existent faiss/run_full_pipeline.py entry point.
The following is a starting point for the original LUMI setup. Cluster modules, project accounts, partitions, model paths, and Slurm resources are site-specific.
module purge
module use /appl/local/csc/modulefiles
module load pytorch/2.5
python3 -m venv --system-site-packages venv
source venv/bin/activate
pip install -r requirements.txtKeep model and dataset caches out of the home directory, for example:
export HF_HOME="/scratch/<your-project>/my_cache"Before submitting a job, update the --account, paths, and resource settings in the relevant run_*.sh file. Descriptor generation is expensive: --num-rewrites=0 is faster, while additional rewrites may improve results. The timings in the original project notes were approximately 10 minutes for 500 documents with zero rewrites and one hour with three rewrites, but they depend on the model, node, and workload.
- The repository targets LUMI, ROCm, and GPU execution. Porting it elsewhere may require changes to the environment and Slurm wrappers.
- The repository relies on a pre-built LUMI AI Factory container image and a virtual environment extension. The container image is not included in the repository. To reproduce the original environment, follow these steps:
- Load modules
module purge
module use /appl/local/laifs/modules
module load lumi-aif-singularity-bindings- Activate container and build venv
export SIF=/appl/local/laifs/containers/lumi-multitorch-u24r64f21m43t29-20260216_093549/lumi-multitorch-full-u24r64f21m43t29-20260216_093549.sif
singularity shell $SIF
Singularity> python -m venv .venv --system-site-packages
Singularity> source .venv/bin/activate
(.venv) Singularity> pip install -r requirements.txtlumi-multitorch-full-u24r64f21m43t29-20260216_093549.sif
- Run scripts using the container and venv in a SLURM job
srun singularity run --rocm --bind /scratch/project_465002530 \
$SIF bash -c "source .venv/bin/activate && python script.py \
--input 'data.jsonl' \
--output 'out.jsonl'
"If you find this repository useful in your research, please cite the following work:
@inproceedings{
tarkkaetal-descriptors,
title={Task-Agnostic Web Document Annotation with LLM-Generated Descriptors},
author={Tarkka, Otto and Henriksson, Erik and Kanerva, Jenna and Ginter, Filip},
year={2026 (forthcoming)},
}