Computer Science editorial
GraphAlignCoder: Aligning Program and Proof Graphs for Code Generation
The core problem
Innovation
GraphAlignCoder consistently outperforms the base model, code-only SFT, and CodeRL across all benchmarks. Compared with CodeRL, it increases the solved count from 38 to 50 on LiveCodeBench v6 and from 16 to 23 on BigCodeBench Hard, corresponding to relative gains of 31.6% and 43.8%. On BigCodeBench Full, it improves from 359 to 363 tasks. These numbers indicate that the largest relative benefit appears on the hardest benchmark, where hidden semantic constraints are most likely to defeat execution-feedback-only training. The consistent direction of improvement across three benchmarks of differing difficulty suggests the gains are not an artifact of a single evaluation set. The reported deltas are summarized below.
| Benchmark | CodeRL | GraphAlignCoder | Relative Gain |
|---|---|---|---|
| LiveCodeBench v6 | 38 | 50 | 31.6% |
| BigCodeBench Hard | 16 | 23 | 43.8% |
| BigCodeBench Full | 359 | 363 | ~1.1% |
The pattern is consistent with the claim that region-level correctness structure provides supervision that binary execution feedback cannot supply.
GraphAlignCoder operates in two stages. First, it constructs an **implementation graph**
Why it matters
Who should read this
Opening member content…