Abstract
Deep hypergraph learning is evaluated almost entirely through leaderboards that rank methods by mean accuracy over a few random seeds, usually without significance testing. Is there a best hypergraph neural network, or does the apparent ordering reflect seed noise? We independently recomputed the node-classification track of DHG-Bench on a single GPU with twenty random seeds (against five upstream) and a different software stack, and applied a four-layer statistical audit to the per-seed accuracies: a reproducibility check, per-dataset paired Wilcoxon tests with Holm correction, an across-datasets Friedman/Iman–Davenport omnibus with Nemenyi and Holm-corrected pairwise tests, and a variance decomposition. Within a single dataset, twenty seeds distinguish most method pairs (74–98%), so the protocol is not underpowered. Across the nine datasets where all 17 methods complete, the omnibus rejects global equality (Kendall’s W=0.45), yet no pair survives Holm correction, and the top methods fall within one critical-difference band. One dataset carries more seed noise than between-method signal and cannot rank methods. The recompute also documents a non-reproducible method, a label-range data fault, and missing per-dataset configurations in the public release. No single method is statistically best across these datasets, so single-leader claims are not supported; we release a reusable significance-aware evaluation protocol.
IPC Classification
Keywords
€ 4.00