A REPRODUCIBLE FRAMEWORK AND STATISTICAL EVALUATION OF NEURAL-GUIDED CONSTRUCTIVE SEARCH

Authors

DOI:

https://doi.org/10.31673/2412-4338.2026.034823

Abstract

Studies of neural combinatorial optimization sometimes assess a trained method using its best run. Such reporting does not distinguish a repeatable improvement from ordinary variation between random starts. This paper describes a software framework and an evaluation protocol designed to make computational experiments repeatable and comparisons statistically interpretable. The illustrative task is constructive folding of protein sequences in the hydrophobic polar model on square and cubic lattices. Six methods share one callable interface. The computational core is separated from an optional web service, and each randomized procedure receives an explicit seed for a local random number generator. The main experiment covers 27 benchmark instances with 20 runs for every instance and method. All methods are assessed using the same definition of the proportion of the known optimum and the success rate. On ten instances, Metropolis Monte Carlo initialized from a trained constructive policy is compared with Monte Carlo initialized from a random conformation using paired seeds and an equal iteration budget. Group means are accompanied by 95% percentile bootstrap intervals, while paired differences are tested with a two sided Wilcoxon signed rank test. In the reported main experiment, Monte Carlo achieved a mean proportion of the optimum of 80.26%, greedy search 63.85%, random selection 30.46%, and the autonomous trained hybrid 25.49%. A trained initial state did not produce a statistically supported improvement over random initialization on any of the ten tested instances; the smallest p value was 0.164. Constructive constraints still preserved the structural validity of completed conformations regardless of policy quality. This distinction between feasible output and solution quality is central to interpreting the result. The practical contribution is a consistent protocol that permits a negative finding about the trained component to be reported without selecting favorable runs. The conclusions are limited to the supplied benchmark set and computational budgets. Ablation results use fewer seeds and should be treated as exploratory; application to other problem classes requires further validation.

Keywords: reproducibility, constructive search, hydrophobic polar model, Wave Function Collapse, bootstrap, Wilcoxon signed rank test, Monte Carlo.

Published

2026-10-01

Issue

Section

Articles