摘要
Although some intelligent large-scale models have been proposed recently, due to the diversity and complexity of urban planning tasks, there are still no domain-specific large-scale models tailored to urban planning that support varied tasks. To address this gap, this paper presents the Semantic Multimodal Analysis and Retrieval approach for Planning (SMART-Plan), an innovative multimodal large model that supports complex and diverse urban planning tasks such as domain knowledge questioning, multimodal data retrieval and image-based plan generation. Another key innovation of this model lies in the automated construction of a domain-specific knowledge graph, which combines textual and visual data to represent urban planning entities and their relationships comprehensively. By leveraging the constructed knowledge graph and the designed three-phase domain fine-tuning, the performance of the model was significantly improved across multiple tasks in urban planning, addressing the challenges of fragmented data and specialized terminology in urban planning. Extensive experiments demonstrated that SMART-Plan significantly outperforms existing models in accuracy, logic and professionalism, with an average improvement of 6.25% and 7.2% in knowledge Q&A and image–text Q&A, compared to state-of-the-art methods.
| 源语言 | 英语 |
|---|---|
| 页(从-至) | 1636-1652 |
| 页数 | 17 |
| 期刊 | Indoor and Built Environment |
| 卷 | 34 |
| 期 | 9 |
| DOI | |
| 出版状态 | 已出版 - 11月 2025 |
联合国可持续发展目标
此成果有助于实现下列可持续发展目标:
-
可持续发展目标 3 良好健康与福祉
学术指纹
探究 'Domain-graph enhanced multi-task multimodal large models for urban planning' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver