WhisperX is an advanced media processing tool designed to transform audio and video content into actionable insights. Leveraging cutting-edge technologies, WhisperX offers unparalleled Summarization, Interactive Question Answering through a sophisticated RAG-based approach, and dynamic Quiz Generation.
- Summarization: Rapidly distill audio and video content into concise, actionable summaries.
- Interactive Question Answering: Utilize an advanced Retrieval-Augmented Generation (RAG) approach to generate precise and contextually relevant answers from transcribed content.
- Quiz Generation: Automatically create engaging quizzes from content summaries to support learning and assessment.
- Llama3 and Gemini: Leading large language models (LLMs) providing robust language processing capabilities.
- OpenAI Whisper/Tiny.en: High-performance transcription tool for converting media to text.
- All-MPNet-Base-V2: Generates high-quality vector embeddings for enhanced semantic understanding.
- FAISS: Implements fast and scalable vector indexing and retrieval for efficient search operations.
WhisperX integrates a sophisticated architecture for seamless media processing:
- Transcription: Convert audio and video to text using Whisper/Tiny.en.
- Embedding Creation: Generate vector embeddings with All-MPNet-Base-V2.
- Indexing and Retrieval: Utilize FAISS for managing and querying embeddings.
- Question Answering: Employ the RAG-based approach to retrieve and generate precise answers.
Ensure you have Ollama installed. Pull the Llama3 model using:
ollama pull llama3-
Clone the Repository
git clone https://github.com/SiddarthAA/whisperx.git cd whisperx -
Install Dependencies And Setup Environment
python setup.py
-
Launch the Application
streamlit run App.py
- Upload: Add your audio or video file via the application interface.
- Select: Choose the appropriate model for summarization, question answering, and quiz generation.
- Configure: Adjust settings including model temperature and language options.
- Generate: Obtain summaries, answers, and quizzes from the uploaded media.
We welcome contributions and collaborations to further enhance WhisperX. If you're interested in collaborating, please get in touch!
- Llama3 and Gemini Teams: For their exceptional language models.
- OpenAI: For Whisper/Tiny.en, which powers our transcription capabilities.
- All-MPNet and FAISS: For their powerful embedding and indexing tools.
- Our Contributors: Heartfelt thanks to everyone who has supported the development of WhisperX.
WhisperX redefines media analysis with its advanced features for summarizing, querying, and quiz generation. Embrace the power of Generative AI to enhance your media processing and learning experiences.
For inquiries or support, please reach out:
- Name: Siddartha Aralakuppe Yogesha
- Email: siddartha_ay@protonmail.com

