(🛠️This is incomplete code currently undergoing testing.)
CFDTM for OpenAlex Paper Collections.
This repository implements a dynamic topic modeling pipeline for OpenAlex paper collections using CFDTM.
The workflow is designed for bibliographic datasets containing metadata such as paper titles, abstracts, publication years, and identifiers.
The main objective is to model how research topics evolve over time based on the textual content of papers.
The input data is a CSV file collected from OpenAlex.
Although the raw dataset may contain multiple metadata fields, the modeling input uses only:
titleabstractpublication_year
The query field is excluded from the modeling input because it reflects the search strategy rather than the actual paper content.
The pipeline consists of the following steps:
- Concatenate the
titleandabstractfields into a unified document text. - Remove duplicate records using
openalex_id,doi, ortitle + year. - Group documents into four-year bins to define sequential time slices.
- Convert the text corpus into bag-of-words representations.
- Train the CFDTM model on the time-sliced corpus.
- Export topic-word distributions over time and document-level dominant topics.
cfdtm-openalex/
├─ README.md
├─ requirements.txt
├─ .gitignore
├─ notebooks/
│ └─ cfdtm_openalex_pipeline.ipynb
├─ src/
│ └─ utils.py
├─ data/
│ ├─ raw/
│ └─ processed/
├─ outputs/
└─ figures/