Hi, thank you for open-sourcing AgentOCR.
I am trying to reproduce the Search results in Table 2 of the paper using the official code and scripts, but my reproduced results are noticeably different from the numbers reported in the paper.
I would like to ask a few questions:
- Were the Table 2 Search results obtained by training with the released
train_search.sh script and then evaluating the resulting checkpoint directly?
- If not, is there any additional training/evaluation code, filtering, checkpoint selection, or hyperparameter setting that was used for the paper results but is not reflected in the released scripts?
- How much variance should be expected across runs / random seeds?
I am wondering whether the gap is mainly due to randomness, or whether I may be using a different evaluation setup from the one used for Table 2.
For reference, my environment is:
- 4 x H200 GPUs
- I used the official
train_search.sh for training
- No intentional algorithmic changes
My reproduced validation results, aligned to the Table 2 dataset order, are:
| NQ |
TriviaQA |
PopQA |
HotpotQA |
2WikiMultiHopQA |
MuSiQue |
Bamboogle |
Avg. |
| 0.359 |
0.555 |
0.369 |
0.327 |
0.301 |
0.118 |
0.332 |
0.346 |
The corresponding per-dataset success_rate values from the same run are:
| NQ |
TriviaQA |
PopQA |
HotpotQA |
2WikiMultiHopQA |
MuSiQue |
Bamboogle |
Avg. |
| 0.373 |
0.594 |
0.420 |
0.358 |
0.337 |
0.126 |
0.169 |
0.340 |
More detailed values from the same run are:
val/nq/test_score = 0.3587
val/triviaqa/test_score = 0.5546
val/popqa/test_score = 0.3694
val/hotpotqa/test_score = 0.3272
val/2wikimultihopqa/test_score = 0.3012
val/musique/test_score = 0.1180
val/bamboogle/test_score = 0.3316
Other relevant metrics from the same run are:
val/success_rate = 0.4056
val/nq_success_rate = 0.3735
val/triviaqa_success_rate = 0.5940
val/popqa_success_rate = 0.4197
val/hotpotqa_success_rate = 0.3578
val/2wikimultihopqa_success_rate = 0.3367
val/musique_success_rate = 0.1256
val/bamboogle_success_rate = 0.1694
val/memory_tokens/mean = 176.0
val/memory_tokens/max = 1026
If helpful, I can also share the full run config / logs.
I would really appreciate any clarification on whether:
- these scripts are exactly the ones used for Table 2,
- or some additional settings are needed for faithful reproduction.
Thank you!
Hi, thank you for open-sourcing AgentOCR.
I am trying to reproduce the Search results in Table 2 of the paper using the official code and scripts, but my reproduced results are noticeably different from the numbers reported in the paper.
I would like to ask a few questions:
train_search.shscript and then evaluating the resulting checkpoint directly?I am wondering whether the gap is mainly due to randomness, or whether I may be using a different evaluation setup from the one used for Table 2.
For reference, my environment is:
train_search.shfor trainingMy reproduced validation results, aligned to the Table 2 dataset order, are:
The corresponding per-dataset
success_ratevalues from the same run are:More detailed values from the same run are:
val/nq/test_score = 0.3587val/triviaqa/test_score = 0.5546val/popqa/test_score = 0.3694val/hotpotqa/test_score = 0.3272val/2wikimultihopqa/test_score = 0.3012val/musique/test_score = 0.1180val/bamboogle/test_score = 0.3316Other relevant metrics from the same run are:
val/success_rate = 0.4056val/nq_success_rate = 0.3735val/triviaqa_success_rate = 0.5940val/popqa_success_rate = 0.4197val/hotpotqa_success_rate = 0.3578val/2wikimultihopqa_success_rate = 0.3367val/musique_success_rate = 0.1256val/bamboogle_success_rate = 0.1694val/memory_tokens/mean = 176.0val/memory_tokens/max = 1026If helpful, I can also share the full run config / logs.
I would really appreciate any clarification on whether:
Thank you!