Experimental topic modeling pipeline developed in the context of a master's thesis. The project combines sentence embeddings, document chunking, dimensionality reduction, soft clustering, embedder fine-tuning, and optional LLM-based topic labeling.
Projekt zawiera eksperymentalny potok modelowania tematycznego. Model dzieli dłuższe dokumenty na fragmenty, tworzy embeddingi, dokonuje redukcji wymiarowości, klastrowania miękkiego, wybiera teksty reprezentatywne i opcjonalnie wykorzystuje LLM do generowania nazw i opisów tematów.
model/
topic_model.py # główny plik metody
ctfidf.py # implementacja c-TF-IDF
loss.py # funkcja straty do finetuningu embeddera
types.py
utils.py
logger.py
clustering/
Clusterer.py # protokół dla algorytmów klastrowania
StudentTKMeans.py # KMeans z miękkimi przypisaniami bazującymi na rozkładzie t-Studenta
llm_clients/
LLMClient.py # protokół klienta LLM
OpenAIClient.py # synchroniczny klient OpenAI
AsyncOpenAIClient.py # asynchroniczny klient OpenAI
python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txtNa Linux/macOS tak:
source .venv/bin/activatefrom sentence_transformers import SentenceTransformer
from model import TopicModel
texts = [
"Text about machine learning and neural networks.",
"Text about elections, parties, and public policy.",
"Text about football, teams, and sports results.",
]
embedding_model = SentenceTransformer("all-MiniLM-L6-v2")
topic_model = TopicModel(
embedding_model=embedding_model,
llm_client=None,
)
probabilities = topic_model.fit_transform(texts)
topics = topic_model.get_topic_info()Jeśli zostanie przekazany klient LLM, model może wygenerować nazwy i opisy tematów.
import os
from sentence_transformers import SentenceTransformer
from llm_clients import OpenAIClient
from model import TopicModel
llm_client = OpenAIClient(
api_key=os.environ["OPENAI_API_KEY"],
model_name="gpt-5.2",
)
topic_model = TopicModel(
embedding_model=SentenceTransformer("all-MiniLM-L6-v2"),
llm_client=llm_client,
dataset_context="Short description of the analyzed dataset.",
)Projekt został przygotowany jako część prac badawczych związanych z pracą magisterską. Celem było zaprojektowanie i zbudowanie metody analizy tematycznej, a nie stworzenie biblioteki produkcyjnej.
This project contains an experimental topic modeling pipeline. The model splits longer documents into chunks, creates embeddings, reduces dimensionality, performs soft clustering, selects representative texts, and optionally uses an LLM to generate topic names and descriptions.
model/
topic_model.py # main method file
ctfidf.py # c-TF-IDF implementation
loss.py # loss function for embedder fine-tuning
types.py
utils.py
logger.py
clustering/
Clusterer.py # protocol for clustering algorithms
StudentTKMeans.py # KMeans with soft assignments based on Student's t-distribution
llm_clients/
LLMClient.py # LLM client protocol
OpenAIClient.py # synchronous OpenAI client
AsyncOpenAIClient.py # asynchronous OpenAI client
python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txtOn Linux/macOS:
source .venv/bin/activatefrom sentence_transformers import SentenceTransformer
from model import TopicModel
texts = [
"Text about machine learning and neural networks.",
"Text about elections, parties, and public policy.",
"Text about football, teams, and sports results.",
]
embedding_model = SentenceTransformer("all-MiniLM-L6-v2")
topic_model = TopicModel(
embedding_model=embedding_model,
llm_client=None,
)
probabilities = topic_model.fit_transform(texts)
topics = topic_model.get_topic_info()When an LLM client is provided, the model can generate topic names and descriptions.
import os
from sentence_transformers import SentenceTransformer
from llm_clients import OpenAIClient
from model import TopicModel
llm_client = OpenAIClient(
api_key=os.environ["OPENAI_API_KEY"],
model_name="gpt-5.2",
)
topic_model = TopicModel(
embedding_model=SentenceTransformer("all-MiniLM-L6-v2"),
llm_client=llm_client,
dataset_context="Short description of the analyzed dataset.",
)The project was prepared as part of research work related to a master's thesis. The goal was to design and build a topic analysis method, not to create a production library.