Skip to content

Question about reproducing Table 2 results with train_search.sh #23

Description

@Fanc0126

Hi, thank you for open-sourcing AgentOCR.

I am trying to reproduce the Search results in Table 2 of the paper using the official code and scripts, but my reproduced results are noticeably different from the numbers reported in the paper.

I would like to ask a few questions:

  1. Were the Table 2 Search results obtained by training with the released train_search.sh script and then evaluating the resulting checkpoint directly?
  2. If not, is there any additional training/evaluation code, filtering, checkpoint selection, or hyperparameter setting that was used for the paper results but is not reflected in the released scripts?
  3. How much variance should be expected across runs / random seeds?

I am wondering whether the gap is mainly due to randomness, or whether I may be using a different evaluation setup from the one used for Table 2.

For reference, my environment is:

  • 4 x H200 GPUs
  • I used the official train_search.sh for training
  • No intentional algorithmic changes

My reproduced validation results, aligned to the Table 2 dataset order, are:

NQ TriviaQA PopQA HotpotQA 2WikiMultiHopQA MuSiQue Bamboogle Avg.
0.359 0.555 0.369 0.327 0.301 0.118 0.332 0.346

The corresponding per-dataset success_rate values from the same run are:

NQ TriviaQA PopQA HotpotQA 2WikiMultiHopQA MuSiQue Bamboogle Avg.
0.373 0.594 0.420 0.358 0.337 0.126 0.169 0.340

More detailed values from the same run are:

  • val/nq/test_score = 0.3587
  • val/triviaqa/test_score = 0.5546
  • val/popqa/test_score = 0.3694
  • val/hotpotqa/test_score = 0.3272
  • val/2wikimultihopqa/test_score = 0.3012
  • val/musique/test_score = 0.1180
  • val/bamboogle/test_score = 0.3316

Other relevant metrics from the same run are:

  • val/success_rate = 0.4056
  • val/nq_success_rate = 0.3735
  • val/triviaqa_success_rate = 0.5940
  • val/popqa_success_rate = 0.4197
  • val/hotpotqa_success_rate = 0.3578
  • val/2wikimultihopqa_success_rate = 0.3367
  • val/musique_success_rate = 0.1256
  • val/bamboogle_success_rate = 0.1694
  • val/memory_tokens/mean = 176.0
  • val/memory_tokens/max = 1026

If helpful, I can also share the full run config / logs.

I would really appreciate any clarification on whether:

  • these scripts are exactly the ones used for Table 2,
  • or some additional settings are needed for faithful reproduction.

Thank you!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions