Repository for A GNN-Guided Predict-and-Search Framework for Mixed-Integer Linear Programming, Han et al., ICLR 2023 (paper, arXiv v4).
The original file structure is retained. CA and IS generation, training, evaluation and plotting are available through the existing scripts. Comments identify paper sections/equations, inherited undocumented choices, and engineering fixes. Generated data, checkpoints and plots are runtime outputs, not additional source files.
Reproduction scope: this implements the released GNN with repaired graph indexing/batching and the paper's published experiment parameters. It does not authenticate the supplied checkpoints or recover the missing original datasets. Exact published numbers cannot be promised: generator settings/seeds and training stopping criteria were not released, and current solver/hardware versions differ. Figure 3 also needs external Neural Diving results; Figure 4 needs the unidentified original WA instance. See the coverage table below.
Linux/WSL, PyTorch, PyG, Ecole, PySCIPOpt, Gurobi, NumPy and matplotlib. Gurobi requires a working license. The paper used PyTorch 1.10.2, Gurobi 9.5.2 and SCIP 8.0.1; the original py38.yaml is retained as the released repository's historical Python 3.8 environment specification.
The prepared, tested Ubuntu-24.04 WSL environment is predict-search-ca (also supports IS):
source ~/anaconda3/etc/profile.d/conda.sh
conda activate predict-search-ca
cd /mnt/c/Users/samer/Desktop/Github/Neural_Diving/Predict-and-Search_MILP_method
python -m pip check
python helper.py self-testTo create the modern environment on another Linux installation:
conda create -y -n predict-search-ca --override-channels -c conda-forge python=3.12 numpy=1.26.4 scipy=1.13.1 ecole=0.8.2 pyscipopt=6.2.1 scip=10.0.3 pip
conda activate predict-search-ca
python -m pip install torch==2.6.0 --index-url https://download.pytorch.org/whl/cu124
python -m pip install numpy==1.26.4 contourpy==1.3.2 torch-geometric==2.6.1 gurobipy==13.0.3 matplotlib==3.10.8
python -m pip checkEcole's conda package is labeled 0.8.2 while its Python metadata reports 0.8.1. NumPy is pinned deliberately. CUDA training works on the installed RTX 4070; use --device cpu if necessary.
gurobi.py: generate LP instances; collect solution pools and bipartite graphs.GCN.py,helper.py: original GNN/features; helper also aggregates/plots results and runs diagnostics.trainPredictModel.py: train and save validation-best/last checkpoints.PredictAndSearch_GRB.py,PredictAndSearch_SCIP.py: baseline, search and fixing evaluations.FixingStrategy_SCIP.py: original fixing entry point, using the shared SCIP evaluation implementation.instance/train/TASK,instance/test/TASK: development and held-out LPs.dataset/TASK/{BG,solution,logs}: collected training data.pretrain/TASK_train,train_logs/TASK_train: new checkpoints and training logs.models: supplied author checkpoints; their training provenance is unavailable.logs/TASK/METHOD,logs/figures: solver outputs, plotted curves and CSV tables.
All commands accept --dataDir to place these same folders under another root. For WSL performance and separate experiments, an empty Linux directory is recommended:
export DATA="$HOME/predict-search-experiments/ca-is-seed0"Section 4 specifies 240 training / 60 validation / 100 test instances. Generate 300 development LPs per family; training makes a seeded 80/20 split. CA and IS are entirely binary. The generator parameters below are explicit assumptions, not recovered paper parameters. Table 3's variable counts alone do not determine them; its printed CA/IS constraint counts are unusual and are not silently swapped.
python gurobi.py --generate --task CA --dataDir "$DATA" --items 300 --bids 1500 --train-count 300 --test-count 100 --seed 0
python gurobi.py --generate --task IS --dataDir "$DATA" --nodes 1500 --graph-type barabasi_albert --affinity 4 --train-count 300 --test-count 100 --seed 0Each split contains generation.json with parameters, per-instance seeds and checksums. Matching re-runs reuse verified LPs; changed settings/data are rejected. Use a new data root for another experiment. Ecole's remaining generator options use its pinned defaults. IS also supports --graph-type erdos_renyi --edge-prob 0.25 as an explicitly different distribution.
Collect feasible solutions (Appendix B: 3,600 seconds per instance, one Gurobi thread):
for TASK in CA IS; do
python gurobi.py --task "$TASK" --dataDir "$DATA" --nWorkers 4 --maxTime 3600 --maxStoredSol 500 --threads 1 --graph-mode corrected
donePoolSearchMode=2, the 500-solution cap and training's first-50 truncation come from the released code; the paper does not specify them. collection.json records settings and hashes. Completed verified instances can be reused. See dataset/dataset.md for formats. IP/WA require externally supplied instances as in the original README; their generator is from ML4CO 2021.
for TASK in CA IS; do
python trainPredictModel.py --task "$TASK" --dataDir "$DATA" --device cuda --seed 0 --epochs 9999 --batch-size 8 --lr 0.003 --max-solutions 50 --graph-mode corrected
doneBatch size 8, Adam and learning rate 0.003 follow Section 4. The 9,999-epoch cap comes from the original code, not a specified paper stopping rule. Best validation loss selects pretrain/TASK_train/model_best.pth; model_last.pth, adjacent config.json and training logs are also saved. Existing checkpoints are protected from overwriting. There is no optimizer-state resume.
The weighted BCE follows Eqs. (3)-(6), with stable softmax/logits. Inherited energy temperatures are 1,000 for CA and 100 for IS; these scales are absent from the paper. --temperature 1 tests the literal formula after converting maximization to minimization. --max-solutions 500 uses all collected solutions. Report these choices and select settings using validation data, never the test set.
--graph-mode corrected repairs the objective-node edge index and zero-range normalization. It keeps the original architecture, extra objective node and unit-edge convention. The paper's edge-coefficient description and half-convolution count do not fully match that architecture. Changing graph conventions requires retraining. CUDA scatter reductions and timed solver paths are not guaranteed bitwise repeatable.
Run the same held-out instances for every method. Evaluation automatically loads the new validation-best checkpoint. Paper Table 4 defaults are CA: Gurobi (600,0,1), SCIP (400,0,10); IS: Gurobi (300,300,20), SCIP (300,300,15).
for TASK in CA IS; do
python PredictAndSearch_GRB.py --task "$TASK" --dataDir "$DATA" --mode bks
python PredictAndSearch_GRB.py --task "$TASK" --dataDir "$DATA" --mode baseline
python PredictAndSearch_GRB.py --task "$TASK" --dataDir "$DATA" --mode search
python PredictAndSearch_SCIP.py --task "$TASK" --dataDir "$DATA" --mode baseline
python PredictAndSearch_SCIP.py --task "$TASK" --dataDir "$DATA" --mode search
python PredictAndSearch_GRB.py --task "$TASK" --dataDir "$DATA" --mode fixing
python FixingStrategy_SCIP.py --task "$TASK" --dataDir "$DATA"
doneBKS uses 3,600 seconds and one Gurobi thread. Other methods use 1,000 seconds and one thread. Gurobi experiments use MIPFocus=1; BKS defaults to 0 because its emphasis is not stated (override --mip-focus explicitly). SCIP uses the upstream aggressive-heuristics setting. Timings are solver time; preprocessing and wall time are recorded separately. A restricted solver's optimal status does not prove optimality of the original MILP.
Each run writes logs/TASK/NAME/results.json and solver logs: settings, data/model hashes, complete-run flag, incumbent trajectories, solutions and original-feasibility checks. Existing run names are protected; --name creates a separate experiment. Small tests may override --time-limit, --limit, --k0, --k1 and --delta.
Figure 5 labels Fixing+Gurobi, whereas Table 4 labels Fixing+SCIP. Both are supported. Figure 5's fixing and search runs must use identical checkpoint, graph mode, seed, k0, k1 and selected partial solution; the plotter verifies this.
To inspect a supplied author checkpoint without retraining, explicitly retain its upstream graph convention:
python PredictAndSearch_GRB.py --task CA --dataDir "$DATA" --checkpoint models/CA.pth --graph-mode upstream --profile upstream --name supplied-ca-upstreamUse --task IS --checkpoint models/IS.pth for IS. --profile upstream reproduces the old neighborhood settings; --profile paper uses Table 4. Upstream graph mode retains the adjacency bug solely for checkpoint comparison, and is not a recommended repaired training mode.
python helper.py plot --dataDir "$DATA" --tasks CA IS --figures 1 2 5This writes PNG/PDF figures, Table 1 CA/IS rows, the Table 2 feature specification with implementation differences, measured Table 3 dimensions, Table 4 specification plus actual run settings, and the plotted numerical values as CSV in logs/figures.
Section 4's metric is abs(OBJ-BKS)/(abs(BKS)+1e-10), averaged per instance, using BKS updated by better compared incumbents. This is not the solver's MIP gap. Curves never use future incumbents; points lacking an incumbent for any test instance are gaps in the curve, with counts reported in CSV. Table 1 gain is (baseline_gap - PS_gap)/baseline_gap * 100, undefined for a zero baseline gap. CA/IS averages are not the published four-family average.
| Paper item | Available through this workflow |
|---|---|
| Figure 1 | Conceptual redraw of the framework; no experimental measurements. |
| Figure 2; Table 1 | CA/IS panels/rows from all four solvers/methods. IP/WA need their own input data and training. |
| Figure 3 | IS/IP plotting with genuine external Neural Diving results; that method is absent from the fork. Simple fixing is not Neural Diving. There is no CA panel in the paper. |
| Figure 4 | Explicit CA/IS/WA/IP perturbation extension below. The original unidentified WA instance and repetition protocol cannot be reconstructed. |
| Figure 5 | CA/IS PS+Gurobi versus fixing with identical partial solutions. |
| Tables 2-4 | Feature/parameter specifications and actual measured dimensions/settings, clearly distinguished. |
For Figure 3, place external results at logs/IS/neural-diving-scip/results.json using the same schema as SCIP results, with config.mode="neural-diving", config.solver="scip" and a nonempty config.provenance identifying the implementation/checkpoint. Match instance hashes, objective sense, budget, completeness and feasible-solution records; then:
python helper.py plot --dataDir "$DATA" --tasks IS --figures 3For an Appendix C extension, first find a certified optimum, select a seeded random partial solution, flip selected bits, and solve the fixed subproblems. Partial-set size, trial count and budgets below are disclosed added choices; timeouts are not counted as proven infeasibility:
python helper.py perturb --dataDir "$DATA" --task CA --instance "$DATA/instance/test/CA/ca_0000.lp" --partial-size 600 --flips 1 2 5 10 20 --trials 100 --seed 0 --reference-time 3600 --time-limit 1000
python helper.py plot --dataDir "$DATA" --tasks CA --figures 4If the reference solve does not prove optimality, the diagnostic stops. A full run can take days: collection alone has a 300 solver-hour cap per family, BKS 100 hours, and each 100-instance evaluation about 27.8 hours. Setup validation uses short, clearly labeled tests; it does not establish the paper's reported gains.
Validation completed in the prepared WSL environment: 18 regression checks; CA/IS generation at 1,500 binary variables; solution collection; one GPU training epoch; and 28 short evaluation solves. All 20 returned incumbents passed original-model feasibility checks; eight restricted IS runs had no incumbent. Figures correctly leave these gaps unreported. These small runs verify the workflow, not full training or the paper's performance claims.