跳到主要导航 跳到搜索 跳到主要内容

LightAlloy-Bench: Benchmarking reasoning abilities of large language models in light alloys

  • Haoyang Xie
  • , Yiran Zhang
  • , Zini Yan
  • , Turab Lookman
  • , Xiangdong Guo
  • , Shuzhou Li
  • , Haoyang Fu
  • , Shiyu Liang
  • , Ziyuan Rao
  • , Xiaoqin Zeng
  • Shanghai Jiao Tong University
  • AiMaterials Research LLC
  • Nanyang Technological University
  • Suzhou Laboratory

科研成果: 期刊稿件文章同行评审

摘要

Light alloys play a critical role in lightweight engineering and sustainable development owing to their high specific strength and resource efficiency. However, their intelligent design remains challenging because of complex structure–process–property relationships and the scarcity of structured, high-quality data. Large language models (LLMs) have recently emerged as a promising paradigm for AI-driven materials research, yet their reasoning capabilities in the context of light alloys have not been systematically assessed. Here, we introduce LightAlloy-Bench, a hierarchical benchmark tailored for light alloy research, comprising a foundational layer with 2300 questions covering domain knowledge and a reasoning layer with 10,096 questions designed to evaluate multi-step reasoning. Using this benchmark, we evaluate 13 general-purpose and 3 materials-domain LLMs. Among all models, GPT-4o achieves the highest reasoning accuracy without any task-specific optimization. Building on this baseline, we further investigate the effects of Chain-of-Thought (CoT) prompting and reinforcement learning (RL)-based optimization. Notably, Qwen3-14B attains the strongest reasoning performance under CoT prompting, outperforming larger-scale LLMs. This result demonstrates that model scale alone does not determine reasoning capability and highlights the critical role of reasoning-oriented optimization strategies. Overall, LightAlloy-Bench establishes a quantitative foundation for advancing LLM-driven reasoning in intelligent light alloy design and offers a generalizable framework for developing reasoning-focused benchmarks across other alloy systems.

源语言英语
文章编号102127
期刊Journal of Magnesium and Alloys
DOI
出版状态已接受/待刊 - 2026
已对外发布

联合国可持续发展目标

此成果有助于实现下列可持续发展目标:

  1. 可持续发展目标 8 - 体面工作和经济增长
    可持续发展目标 8 体面工作和经济增长
  2. 可持续发展目标 12 - 负责任消费和生产
    可持续发展目标 12 负责任消费和生产

学术指纹

探究 'LightAlloy-Bench: Benchmarking reasoning abilities of large language models in light alloys' 的科研主题。它们共同构成独一无二的指纹。

引用此