跳到主要导航 跳到搜索 跳到主要内容

Warm-start or cold-start? A comparison of generalizability in gradient-based hyperparameter tuning

  • Yubo Zhou
  • , Jun Shu
  • , Chengli Tan
  • , Haishan Ye
  • , Quanziang Wang
  • , Junmin Liu
  • , Deyu Meng
  • , Ivor Tsang
  • , Guang Dai
  • Xi'an Jiaotong University
  • State Grid Corporation of China
  • Northwestern Polytechnical University Xian
  • Nanyang Technological University

科研成果: 期刊稿件文章同行评审

摘要

Bilevel optimization (BO) has garnered increasing attention in hyperparameter tuning. BO methods are commonly employed with two distinct strategies for the inner-level: cold-start, which uses a fixed initialization, and warm-start, which uses the last inner approximation solution as the starting point for the inner solver each time, respectively. Previous studies mainly stated that warm-start exhibits better convergence properties, while we provide a detailed comparison of these two strategies from a generalization perspective. Our findings indicate that, compared to the cold-start strategy, warm-start strategy exhibits worse generalization performance, such as more severe overfitting on the validation set. To explain this, we establish generalization bounds for the two strategies. We reveal that warm-start strategy produces a worse generalization upper bound due to its closer interaction with the inner-level dynamics, naturally leading to poor generalization performance. Inspired by the theoretical results, we propose several approaches to enhance the generalization capability of warm-start strategy and narrow its gap with cold-start, especially a novel random perturbation initialization method. Experiments validate the soundness of our theoretical analysis and the effectiveness of the proposed approaches.

源语言英语
文章编号108647
期刊Neural Networks
199
DOI
出版状态已出版 - 7月 2026

学术指纹

探究 'Warm-start or cold-start? A comparison of generalizability in gradient-based hyperparameter tuning' 的科研主题。它们共同构成独一无二的指纹。

引用此