Manusya is a personal AI assistant you can talk or type to. It listens (or reads what you type), figures out what topic you're talking about (like physics, cooking, or finance), and remembers that topic separately from every other topic — almost like it keeps a different notebook for each subject. Over time it learns more about each topic from your conversations, so its answers on that subject should get better the more you use it.
Everything it saves is locked (encrypted) so nobody else can read it, and if someone copies the program to a different computer, all the saved data becomes unreadable and the program starts fresh — nothing leaks.
- Listen and speak: Click the microphone button and talk to it right in your browser, or just type instead. It turns your speech into text and can read its answers out loud (either automatically, or only when you tap "Listen" on a reply). It understands/speaks English, Hindi, Hinglish (mixed Hindi-English), Spanish, French, German, Italian, Dutch, Arabic, Russian, Chinese, Japanese, Vietnamese, Indonesian, Swahili, Turkish, and Polish.
- Clean up messy speech: If what you said (or what speech-recognition heard) comes out a bit garbled or badly phrased, a small AI model tidies it up before answering, and does the same to make the final answer sound natural when read aloud.
- Sort conversations into topics automatically: Nobody has to tell it "this is about cooking" —
it figures that out itself by comparing what you say to topics it has seen before. The same
topic always gets grouped together consistently. Topic names, however, are deliberately
meaningless — random codes like
domain_4987c8709arather than words likecooking— on purpose: since these names double as file/folder names and show up in the dashboard, using the real topic as the name would let anyone who glances at your files or screen read off exactly what subjects you've been asking about. The grouping is smart; the label is intentionally blank. You can also cap how many topics it's allowed to create (see "Domain" in the dashboard below); once that limit is hit, anything new gets folded into one shared "general" topic instead of endlessly multiplying. - A dedicated mini-assistant per topic: Each topic gets its own private copy of the AI model and its own notes folder. It answers using what it already knows plus anything relevant you've told it before on that topic (this is called "RAG" — basically, it looks up your own past notes before answering). Periodically (or when you ask it to), it can actually retrain that topic's model on everything it has collected, so it keeps improving.
- New topics just work: The moment you bring up a brand-new subject, a brand-new topic-notebook and mini-assistant are created for it automatically — no setup needed.
- Locked data: All your saved conversations, notes, and topic information are encrypted. The AI model files themselves are left as-is (they need to be loadable by the AI software directly). There's also a "reset" button that wipes everything and starts over with fresh locks.
- A simple on-screen dashboard: A web page where you can talk/type to the assistant. The
message box accepts multi-line messages (Shift+Enter for a new line, Enter to send) and can be
resized by dragging its bottom-right corner, and the conversation area expands to use the rest
of the window. The sidebar settings are grouped into collapsible sections so only what you're
changing is on screen at once:
- Speech: which language you speak in, which language answers come back in, whether to read replies aloud automatically, and how accurate (vs. fast) speech recognition should be.
- Data Handling: turn on/off saving of your conversation, delete your saved conversation, or export it as a small text file you could use to train an AI model later.
- Processing Mode: choose whether Manusya runs its AI models on the GPU, the CPU, or Auto (the default — prefers a GPU if your computer has one, otherwise uses the CPU). It tells you right there whether it actually found a graphics card on your machine. This same panel also shows how many processing "threads" your computer has available (e.g. "16 threads detected on this machine") and lets you choose how many of them Manusya is allowed to use — this only matters while it's actually running on the CPU (either because you chose "CPU", or because "Auto"/"GPU" fell back to it), and more threads can mean faster replies on a multi-core computer. "Auto" for threads lets the AI software decide for itself.
- Model: which AI model is used for the general text clean-up step and which one is cloned into new topics, and how many words ("tokens") each is allowed to generate per reply. Either pick an id already downloaded, type a new Hugging Face model id (downloaded automatically the first time it's needed, same as the very first run of the whole program), or click "Browse..." to pick a model folder straight from your own computer.
- Domain: shows how many topics Manusya has learned so far, and lets you cap that number -- once the cap is reached, anything new is folded into a shared "general" topic (with its own model and notes) instead of starting another dedicated one. -1 means no limit.
- Remembers your settings: If you close the dashboard and reopen it later, it remembers the language you had picked, whether saving was on, and so on — you don't have to set it up again.
You need Python installed, plus the espeak-ng system package (used to read replies aloud) --
e.g. sudo apt install espeak-ng on Debian/Ubuntu, brew install espeak-ng on macOS. Then, in a
terminal, inside this folder:
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtThe AI models it uses are small and free, and get downloaded automatically the first time you run the program (this can take a few minutes depending on your internet). Everything runs on an ordinary computer — no special graphics card needed — though answers may take a little while to generate since it's all running on a normal CPU.
# Open the dashboard (the main way to use it)
python app.pyThen open http://127.0.0.1:8000 in your browser. Click the 🎤 button to talk (your browser will ask for microphone permission the first time), or just type in the message box instead.
The dashboard is a plain HTML/CSS/JS page talking to a small local web server (no separate frontend build step needed), which also makes it straightforward to later deploy this same server to a real machine rather than only running it on your own computer.
There are also two check-up scripts that make sure everything is working correctly behind the scenes (useful after installing or updating things, not needed for everyday use):
python scripts/smoke_test.py # checks the "understand topics and learn" part
python scripts/verify_speech_and_vault.py # checks speech, translation, and the save/lock featureEvery adjustable setting (which AI models to use, which languages are supported, how "different"
a topic needs to be before it's treated as a new one, how often to retrain, where files are saved,
etc.) lives in one place: config/config.yaml. You can open it in any text editor to see or change
these.
Every instruction Manusya gives its AI models -- what tone to use, what job to do, how to answer --
lives in its own plain text file in the prompts/ folder, not buried in code. Open any of these in
a text editor, change the wording, save, and restart Manusya to see the new behavior; no programming
knowledge needed. Each one is used exactly where its name suggests:
input_cleanup_system.txt-- Instructions for tidying up what you said (or what speech recognition heard) into one clear, grammatically-correct sentence before Manusya thinks about an answer. Edit this if cleaned-up text is coming out too short, too long, or changing your meaning.output_cleanup_system.txt-- Instructions for rewriting the AI's raw answer into something that sounds natural when read aloud. Edit this if replies sound too robotic, too wordy, or too short.topic_label_system.txt-- A one-line instruction telling the AI "you are a subject classifier." Rarely needs changing.topic_label_user.txt-- The actual question asked to figure out which broad subject (physics, cooking, finance, ...) a message belongs to -- this is what decides which topic/domain a conversation gets filed under. Contains a{text}placeholder that gets replaced with the message being classified; keep that placeholder if you edit the wording around it. Edit this if topics are being classified too broadly (everything lumped together) or too narrowly (near- identical questions ending up in different topics) -- thoughdissimilarity_thresholdinconfig/config.yamlis usually the better first thing to tune for that.domain_answer_system_with_context.txt-- Instructions used when answering a question and Manusya has found relevant notes from earlier conversations on that same topic. Contains a{context}placeholder that gets replaced with those notes; keep that placeholder if you edit the wording around it.domain_answer_system_no_context.txt-- The same, simpler instruction used when there are no relevant earlier notes yet (e.g. the very first question on a brand-new topic).
A quick way to think about it: the first two files control how things sound (input/output
polish), topic_label_* controls where a conversation gets filed, and domain_answer_* controls
how the actual answer gets written.
If the prompts/ folder or any individual file inside it ever goes missing (accidentally deleted,
an incomplete copy of the project, etc.), Manusya doesn't break -- it quietly falls back to the
same wording built into the code, one file at a time, so a single missing file doesn't affect the
others.
- Everything the assistant remembers (topics, chat history, notes, and your dashboard settings) is locked with a secret key that is tied to this specific computer. If the program (or its data) is copied to another computer, the lock won't open there — the data is unreadable, and the program simply starts completely fresh instead of showing an error. The same thing happens if the settings file is ever missing or empty: Manusya treats that as "nothing to trust" and starts fresh too, rather than continuing with mismatched data.
- The dashboard's own "save conversation" feature uses a separate lock based on a passphrase you choose yourself — so even on the same computer, that saved file needs your passphrase to open, and moving it to a different computer (or a full reset) doesn't affect it at all.
- The AI model files are not locked (they need to be readable by the AI software to work), but they don't contain your personal conversations — those are stored and locked separately.
- A "reset" option is available to wipe everything and generate brand-new locks, if you ever want a completely clean slate.
- The "how different does a topic need to be before it counts as a new one" setting was tuned by
testing it against real examples rather than left at an arbitrary guess — see the note next to
dissimilarity_thresholdinconfig/config.yamlfor what was tested. - Recognizing synthetic, non-English speech (e.g. reading a computer-generated Spanish/Hindi
reply back through speech recognition, as in testing) is noticeably rougher than English --
real human speech in those languages is expected to fare better than a robotic voice reading the
same text. This is a quality ceiling of the small ("tiny") speech-recognition model configured by
default, not a pipeline bug. A larger
speech.stt.model_sizeinconfig/config.yaml(e.g. "small" or "base") should help if you hit this in practice. - Because the AI model used here is small and quick (so it can run on an ordinary computer), its answers can occasionally mix in unrelated information once it starts pulling from your past notes — this is a limitation of using a small, fast model rather than a bug in how the assistant is put together.