collection of text2cypher datasets, evaluations, and finetuning instructions
-
Updated
Jun 13, 2024 - Jupyter Notebook
collection of text2cypher datasets, evaluations, and finetuning instructions
Repository for organizing datasets and papers used in Open LLM.
SyGra - Graph-oriented Synthetic data generation Pipeline
A data-centric AI package for ML/AI. Get the best high-quality data for the best results. Discord: https://discord.gg/t6ADqBKrdZ
Better AI Predictions ⚡
A collection of recent open-source math datasets for training and evaluating Math LLMs
A framework to analyze how AGI/ASI might emerge from decentralized, adaptive systems, rather than as the fruit of a single model deployment. It also aims to present orientation as a dynamic and self-evolving Magna Carta, helping to guide the emergence of such phenomena.
[ICML 2026] CapBencher toolkit: Give your LLM benchmark a built-in alarm for leakage and gaming
Efficiently fetch and perform sentiment analysis (Turkish Only) on eksisozluk.com entries using Rust
WikiText syntax dataset generation pipeline and open dataset for auto UI generation in TiddlyWiki. (WIP)
Synthetically Generating Intent-Aware Information-Seeking Dialogues! Useful for various tasks such as training/evaluating User Intent Predictors with the possibility to training/evaluating on real human dialogues. The backbone LLM of SOLID is Zephyr-7b-beta.
🔍 High-throughput AST data mining and semantic code sifting engine for AI datasets and automated codebase transformations.
LLM-Powered Dataset Creation Tool
Convert multi-speaker audio files to structured chat data for LLMs
A collection of Persian poems structured for NLP and LLM tasks. Each poem is stored as a separate file, organized by poet, and formatted for easy use in training, fine-tuning, or text analysis workflows.
Sievio turns GitHub, local repos, and web PDFs into clean JSONL for LLM pretraining, fine-tuning, and RAG. It offers structure-aware chunking, reliable Unicode decoding, pluggable QC and safety checks, plus optional dataset cards and deduplication.
Stratified LLM Subsets delivers diverse training data at 100K-1M scales across pre-training (FineWeb-Edu, Proof-Pile-2), instruction-following (Tulu-3, Orca AgentInstruct), and reasoning distillation (Llama-Nemotron). Embedding-based k-means clustering ensures maximum diversity across 5 high-quality open datasets.
PARROT (Performance Assessment of Reasoning and Responses On Trivia) is a novel benchmarking framework designed to evaluate Large Language Models (LLMs) on real-world, complex, and ambiguous QA tasks.
A modified dataset consisting of English dialogs between a user and an assistant discussing movie preferences in natural language.
A high-performance, asynchronous Go web crawler built to extract LLM-ready Markdown from any website.
To associate your repository with the llm-datasets topic, visit your repo's landing page and select "manage topics."