Archive/Is There a Best Hypergraph Neural Network? A Significance-Aware Recomputation and Statistical Audit of DHG-Bench
Is There a Best Hypergraph Neural Network? A Significance-Aware Recomputation and Statistical Audit of DHG-Bench
Valeriya V. Tynchenko, Sergei O. Kurashkin, Aleksei S. Borodulin et al.
July 31, 2026
en

Abstract

Deep hypergraph learning is evaluated almost entirely through leaderboards that rank methods by mean accuracy over a few random seeds, usually without significance testing. Is there a best hypergraph neural network, or does the apparent ordering reflect seed noise? We independently recomputed the node-classification track of DHG-Bench on a single GPU with twenty random seeds (against five upstream) and a different software stack, and applied a four-layer statistical audit to the per-seed accuracies: a reproducibility check, per-dataset paired Wilcoxon tests with Holm correction, an across-datasets Friedman/Iman–Davenport omnibus with Nemenyi and Holm-corrected pairwise tests, and a variance decomposition. Within a single dataset, twenty seeds distinguish most method pairs (74–98%), so the protocol is not underpowered. Across the nine datasets where all 17 methods complete, the omnibus rejects global equality (Kendall’s W=0.45), yet no pair survives Holm correction, and the top methods fall within one critical-difference band. One dataset carries more seed noise than between-method signal and cannot rank methods. The recompute also documents a non-reproducible method, a label-range data fault, and missing per-dataset configurations in the public release. No single method is statistically best across these datasets, so single-leader claims are not supported; we release a reusable significance-aware evaluation protocol.

IPC Classification

G06H04H01

Keywords

therebesthypergraphneuralnetworksignificance-awarerecomputationstatisticalauditdhg-benchmachinelearningknowledgeextractiondeepevaluatedalmostentirelythroughleaderboardsrankmeanaccuracyrandom
Reference this publication

€ 4.00