Skip to main navigation Skip to search Skip to main content

Towards multi-language repository-level code generation: From-scratch to guided tasks

  • Jingjing Liu
  • , Silin Li
  • , Zeming Liu*
  • , Zihao Cheng
  • , Yuhang Guo
  • , Yuanfang Guo
  • , Yunhong Wang
  • , Haifeng Wang
  • *Corresponding author for this work
  • Beihang University
  • Beijing Institute of Technology
  • Baidu Inc

Research output: Contribution to journalArticlepeer-review

Abstract

Repository-level code generation constitutes a fundamental building block for automated software development. Consequently, numerous benchmarks have been proposed to evaluate the capabilities of large language models (LLMs) in this domain. However, existing benchmarks are largely limited to a single programming language and a fixed granularity level. To address these two challenges, we introduce ReCode-bench, a multi-language benchmark covering 7 widely-used programming languages and comprising three repository-level code generation tasks. These tasks are designed around varying levels of requirement granularity and include full project creation from scratch as well as guided development based on structural or functional specifications. In the latter, intentionally introduced requirements are treated as positive noise to better reflect realistic development scenarios. To enhance LLM robustness across these tasks, we propose RepoGenesis, a GRPO-based reinforcement learning framework that incorporates 3 distinct reward signals: structural similarity to human-written repositories, syntactic correctness verified through abstract syntax tree analysis, and functional validity confirmed via unit test execution. We evaluated 8 LLMs on ReCode-bench and found that even Claude-Sonnet-4—currently among the strongest code generation models—achieves an average pass@1 score of less than 4% across the three tasks. However, after training with RepoGenesis, Qwen2.5-coder-7B-Instruct attains performance comparable to Claude-Sonnet-4 (>100B).

Original languageEnglish
Article number133204
JournalNeurocomputing
Volume679
DOIs
Publication statusPublished - 28 May 2026
Externally publishedYes

Keywords

  • GRPO
  • Large language model
  • Multi-language
  • RepoGenesis
  • Repository-level code generation

Fingerprint

Dive into the research topics of 'Towards multi-language repository-level code generation: From-scratch to guided tasks'. Together they form a unique fingerprint.

Cite this