Skip to content
dawid-walkiewiczPublic

About

Topic modeling method combining sentence embeddings, document chunking, dimensionality reduction, soft clustering, PEFT, and LLM-based topic labeling.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

26 Commits

Folders and files

Repository files navigation

Topic Modeling Pipeline

Experimental topic modeling pipeline developed in the context of a master's thesis. The project combines sentence embeddings, document chunking, dimensionality reduction, soft clustering, embedder fine-tuning, and optional LLM-based topic labeling.

Languages: Polski | English

Wersja Polska

Opis

Projekt zawiera eksperymentalny potok modelowania tematycznego. Model dzieli dłuższe dokumenty na fragmenty, tworzy embeddingi, dokonuje redukcji wymiarowości, klastrowania miękkiego, wybiera teksty reprezentatywne i opcjonalnie wykorzystuje LLM do generowania nazw i opisów tematów.

Struktura repozytorium

model/
  topic_model.py      # główny plik metody
  ctfidf.py           # implementacja c-TF-IDF
  loss.py             # funkcja straty do finetuningu embeddera
  types.py
  utils.py
  logger.py

clustering/
  Clusterer.py        # protokół dla algorytmów klastrowania
  StudentTKMeans.py   # KMeans z miękkimi przypisaniami bazującymi na rozkładzie t-Studenta

llm_clients/
  LLMClient.py        # protokół klienta LLM
  OpenAIClient.py     # synchroniczny klient OpenAI
  AsyncOpenAIClient.py # asynchroniczny klient OpenAI

Instalacja

python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txt

Na Linux/macOS tak:

source .venv/bin/activate

Podstawowe użycie

from sentence_transformers import SentenceTransformer

from model import TopicModel

texts = [
    "Text about machine learning and neural networks.",
    "Text about elections, parties, and public policy.",
    "Text about football, teams, and sports results.",
]

embedding_model = SentenceTransformer("all-MiniLM-L6-v2")

topic_model = TopicModel(
    embedding_model=embedding_model,
    llm_client=None,
)

probabilities = topic_model.fit_transform(texts)
topics = topic_model.get_topic_info()

Etykietowanie przez LLM

Jeśli zostanie przekazany klient LLM, model może wygenerować nazwy i opisy tematów.

import os

from sentence_transformers import SentenceTransformer

from llm_clients import OpenAIClient
from model import TopicModel

llm_client = OpenAIClient(
    api_key=os.environ["OPENAI_API_KEY"],
    model_name="gpt-5.2",
)

topic_model = TopicModel(
    embedding_model=SentenceTransformer("all-MiniLM-L6-v2"),
    llm_client=llm_client,
    dataset_context="Short description of the analyzed dataset.",
)

Kontekst akademicki

Projekt został przygotowany jako część prac badawczych związanych z pracą magisterską. Celem było zaprojektowanie i zbudowanie metody analizy tematycznej, a nie stworzenie biblioteki produkcyjnej.

English Version

Overview

This project contains an experimental topic modeling pipeline. The model splits longer documents into chunks, creates embeddings, reduces dimensionality, performs soft clustering, selects representative texts, and optionally uses an LLM to generate topic names and descriptions.

Repository Structure

model/
  topic_model.py      # main method file
  ctfidf.py           # c-TF-IDF implementation
  loss.py             # loss function for embedder fine-tuning
  types.py
  utils.py
  logger.py

clustering/
  Clusterer.py        # protocol for clustering algorithms
  StudentTKMeans.py   # KMeans with soft assignments based on Student's t-distribution

llm_clients/
  LLMClient.py        # LLM client protocol
  OpenAIClient.py     # synchronous OpenAI client
  AsyncOpenAIClient.py # asynchronous OpenAI client

Installation

python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txt

On Linux/macOS:

source .venv/bin/activate

Basic Usage

from sentence_transformers import SentenceTransformer

from model import TopicModel

texts = [
    "Text about machine learning and neural networks.",
    "Text about elections, parties, and public policy.",
    "Text about football, teams, and sports results.",
]

embedding_model = SentenceTransformer("all-MiniLM-L6-v2")

topic_model = TopicModel(
    embedding_model=embedding_model,
    llm_client=None,
)

probabilities = topic_model.fit_transform(texts)
topics = topic_model.get_topic_info()

LLM Labeling

When an LLM client is provided, the model can generate topic names and descriptions.

import os

from sentence_transformers import SentenceTransformer

from llm_clients import OpenAIClient
from model import TopicModel

llm_client = OpenAIClient(
    api_key=os.environ["OPENAI_API_KEY"],
    model_name="gpt-5.2",
)

topic_model = TopicModel(
    embedding_model=SentenceTransformer("all-MiniLM-L6-v2"),
    llm_client=llm_client,
    dataset_context="Short description of the analyzed dataset.",
)

Academic Context

The project was prepared as part of research work related to a master's thesis. The goal was to design and build a topic analysis method, not to create a production library.

About

Topic modeling method combining sentence embeddings, document chunking, dimensionality reduction, soft clustering, PEFT, and LLM-based topic labeling.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages