Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 

Repository files navigation

CFDTM-OpenAlex

(🛠️This is incomplete code currently undergoing testing.)

CFDTM for OpenAlex Paper Collections.

Overview

This repository implements a dynamic topic modeling pipeline for OpenAlex paper collections using CFDTM.
The workflow is designed for bibliographic datasets containing metadata such as paper titles, abstracts, publication years, and identifiers.

The main objective is to model how research topics evolve over time based on the textual content of papers.

Data

The input data is a CSV file collected from OpenAlex.
Although the raw dataset may contain multiple metadata fields, the modeling input uses only:

  • title
  • abstract
  • publication_year

The query field is excluded from the modeling input because it reflects the search strategy rather than the actual paper content.

Method

The pipeline consists of the following steps:

  1. Concatenate the title and abstract fields into a unified document text.
  2. Remove duplicate records using openalex_id, doi, or title + year.
  3. Group documents into four-year bins to define sequential time slices.
  4. Convert the text corpus into bag-of-words representations.
  5. Train the CFDTM model on the time-sliced corpus.
  6. Export topic-word distributions over time and document-level dominant topics.

Repository Structure

cfdtm-openalex/
├─ README.md
├─ requirements.txt
├─ .gitignore
├─ notebooks/
│  └─ cfdtm_openalex_pipeline.ipynb
├─ src/
│  └─ utils.py
├─ data/
│  ├─ raw/
│  └─ processed/
├─ outputs/
└─ figures/

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors