A multi-agent system (MAS) that performs an automated literature review: it turns a natural-language research question into arXiv search queries, filters the results for relevance, downloads and stores the papers, extracts methodology/findings/future work from each one, and synthesises a final report on common themes and research gaps.
Built with SPADE (XMPP-based agent platform) and SPADE-BDI (AgentSpeak/Jason reasoning), using the Gemini API for language understanding, the arXiv API for retrieval, and the Jina Reader API for full-text extraction.
Course: TIES454 — Agentic Technologies for Developers, University of Jyväskylä Team G.O.A.T: Abdelaziz Ibrahim, Besher Alkurdi, Bishwash Khanal
Manual literature review is slow and error-prone. A PhD student entering a new field, or a team looking for contradictory findings across studies, has to deal with:
- Information overload — the volume of published papers is overwhelming
- Quality assessment — judging relevance and reliability of each source is difficult
- Synthesis complexity — connecting findings across heterogeneous studies requires expertise
- Time constraints — extracting meaningful insight from every source takes too long
The system addresses this through two scenarios, mirroring the design documents:
- Comprehensive literature collection — build a knowledge base of relevant papers for a question.
- Content analysis and synthesis — read the collected papers and produce a synthesis report.
Six agents communicate over XMPP with JSON payloads. Every message carries a type metadata field
(see models.py, class MessageType), and each agent installs a SPADE Template that filters for the
message type it consumes.
| Agent | Role | Behaviour (src/) |
Behaviour (bdi_agent/) |
Module |
|---|---|---|---|---|
| Query Construction | QueryFormulator |
OneShotBehaviour |
BDI (query_construction.asl) |
agents/query.py / agents/query_bdi.py |
| Search | SearchCoordinator |
CyclicBehaviour |
CyclicBehaviour |
agents/search.py |
| Relevant | ResultValidator |
CyclicBehaviour |
BDI (relevant.asl) |
agents/relevant.py / agents/relevant_bdi.py |
| Knowledge Aggregator | ContentCollector |
OneShotBehaviour |
BDI (knowledge_aggregator.asl) |
agents/knowledge.py / agents/knowledge_bdi.py |
| Analysis | AnalyzePapers |
OneShotBehaviour |
OneShotBehaviour |
agents/analysis.py |
| Synthesis | SynthesizeReport |
OneShotBehaviour |
OneShotBehaviour |
agents/synthesis.py |
sequenceDiagram
participant U as User (UI / main.py)
participant Q as QueryConstruction
participant S as Search
participant R as Relevant
participant K as KnowledgeAggregator
participant A as Analysis
participant Y as Synthesis
U->>Q: research_query
Q->>S: search_params (3 arXiv queries via Gemini)
S->>R: search_results (arXiv API, parsed XML)
R->>Q: refined_query (if too few relevant papers)
Q->>S: search_params (refined)
R->>K: relevant_papers (score >= threshold)
K->>A: knowledge_ready (folder_path)
A->>Y: analysis_ready (results_path)
Y-->>U: final_report.json
What each agent does
- Query Construction — asks Gemini to turn the question into three arXiv-syntax queries with explanations; also handles refinement requests, re-prompting with the previous result IDs.
- Search — runs every query against
export.arxiv.org/api/query, parses the Atom XML into structured records (id, title, summary, authors, published, pdf/page URLs, categories). - Relevant — scores each paper against the research question with Gemini, keeps papers above
relevance_threshold, and either asks for a refined query (when the LLM flags refinement and fewer than 5 papers passed) or forwards the set downstream. - Knowledge Aggregator — deduplicates papers, resolves an HTML version of each one
(
arxiv.org/html/..., falling back toar5iv.org), fetches the full text as Markdown through the Jina Reader API, and writes the knowledge base to disk (top 10 papers). - Analysis — reads each stored Markdown paper and extracts methodology, findings, and future work as structured JSON.
- Synthesis — consumes the aggregated analysis and produces common themes, research gaps, and suggested future work as the final report.
The repository keeps both milestone implementations side by side; they share the same message protocol, services, and output format.
src/— plain SPADE. All agents are proceduralOneShot/Cyclicbehaviours. Streamlit UI.bdi_agent/— the BDI version. Query Construction, Relevant, and Knowledge Aggregator areBDIAgents whose reasoning lives in AgentSpeak.aslplans underbdi_agent/asl/, with Python custom actions (e.g..create_search_queries,.send_search_params) bridging beliefs to the LLM and XMPP layers. Gradio UI. This is the version described inreport_2/.
├── src/ # SPADE implementation (report_1)
│ ├── main.py # Boots all six agents, sends a hard-coded question
│ ├── ui.py # Streamlit front-end
│ ├── agents/ # One module per agent
│ ├── services/ # arXiv + Gemini API clients
│ ├── config.py # CONFIG dict loaded from .env
│ └── models.py # MessageType constants
├── bdi_agent/ # SPADE-BDI implementation (report_2)
│ ├── main.py
│ ├── ui_gradio.py # Gradio front-end
│ ├── agents/ # *_bdi.py modules + shared search/analysis/synthesis
│ └── asl/ # AgentSpeak plan libraries
├── proposal/ # Project proposal (LaTeX + PDF)
├── design/ # GAIA design document, diagrams, Quarto/reveal.js slides
├── report_1/ # First implementation report (SPADE agents)
├── report_2/ # Final report (BDI agents, both scenarios)
└── .github/ # Copilot instructions used during development
- Python 3.10+
- A Gemini API key
- A Jina Reader API key
Pick an implementation and install its dependencies (the BDI version additionally needs spade-bdi,
and its Gradio UI needs gradio, which is not in requirements.txt):
cd bdi_agent # or: cd src
python -m venv venv
source venv/bin/activate
pip install -r requirements.txt
pip install gradio # only for bdi_agent/ui_gradio.pycp .env.example .envGEMINI_API_KEY="your-key"
JINA_API_KEY="your-key"Other knobs live in config.py:
| Key | Default | Meaning |
|---|---|---|
timeout |
60 |
Seconds a behaviour waits for an incoming message (doubled for the aggregator/analysis/synthesis stages) |
max_results |
20 |
Papers requested from arXiv per query |
relevance_threshold |
0.7 |
Relevance cut-off; scaled to the LLM's 0–10 scale internally |
The agents use JIDs on localhost, so a local XMPP server must be running first. SPADE 4 ships with an
embedded pyjabber server; start it in a separate terminal (it
creates/uses server.db):
spade runThen, from inside src/ or bdi_agent/:
python main.py # headless run with the question hard-coded in main.py
streamlit run ui.py # src/ — web UI
python ui_gradio.py # bdi_agent/ — web UIAgents register themselves on first connect with the password password. main.py keeps the system
alive for 5 minutes and then stops every agent; the UIs use a shorter window and then display the
results.
Note:
src/ui.pyonly starts the first four agents (scenario 1) and reads the newest JSON file fromsrc/results/, which is where an earlier version of the aggregator wrote its output — current runs write toknowledge_bases/instead.bdi_agent/ui_gradio.pyruns all six agents and browsesknowledge_bases/, so it reflects the full pipeline.
Each run writes a timestamped directory named after the research question:
knowledge_bases/<slugified_question>_<YYYYmmdd_HHMMSS>/
├── 2104.00746v1.md # full text of each paper, via Jina Reader
├── ...
├── research.json # knowledge base: question, papers, scores, URLs, timestamp
├── analysis.json # per-paper methodology / findings / future work
└── final_report.json # common_themes, research_gaps, suggested_future_work
research.json has the shape:
{
"research_question": "",
"papers": [
{ "id": "", "title": "", "abstract": "", "authors": [], "relevance_score": 0.0, "url": "" }
],
"timestamp": ""
}Sample outputs from real runs are checked in under src/knowledge_bases/ and
bdi_agent/knowledge_bases/.
| Document | Source | Contents |
|---|---|---|
| Proposal | proposal/main.tex → main.pdf |
Problem domain, importance, challenges, the two scenarios |
| Design | design/main.tex → main.pdf |
GAIA analysis: environmental, role, interaction, agent, and service models |
| Slides | design/index.qmd → presentation.pdf |
reveal.js deck of the design (built with Quarto) |
| Report 1 | report_1/report.tex → report.pdf |
First four agents, roles/behaviours, sequence diagram, console screenshots |
| Report 2 | report_2/report.tex → report.pdf |
Final report: BDI agents with goals, both scenarios, block diagram, screenshots |
Diagrams live in design/images/ (PNG and SVG).
Build the LaTeX documents with latexmk -pdf main.tex (or report.tex), and the slides with
quarto render design/index.qmd.
| Milestone | Date |
|---|---|
| Project domain agreement | 08.04.2025 |
| MAS design using an AOSE (GAIA) methodology | 24.04.2025 |
| First agent(s) implemented with SPADE | 06.05.2025 |
| Full MAS implemented with a MAS framework | 20.05.2025 & 22.05.2025 |