Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Computer Science editorial

Open AccessOA2026

Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models

CLIC: A Visual Analytics Approach to Characterize LLM Coding Behavior via Token-Frequency Analysis
Junpeng Wang; Yuzhong Chen; Menghai Pan; Uday Singh Saini; Yiwei Caiยท 2026ยท DOI 10.48550/arXiv.2609.22097

The core problem

The evaluation of large language models (LLMs) on coding tasks has primarily focused on performance metrics such as pass@k. As LLMs continue to advance, many models now meet baseline performance requirements, reducing the discriminative power of performance-based evaluation alone. Yet a key question remains largely unexplored: how do LLMs differ in their coding behavior? This paper addresses this gap by proposing CLIC (Code Learning for Identification and Comparison), a visual analytics approach that characterizes LLM coding behavior through token-frequency analysis. The authors argue that understanding behavioral differences is crucial for tasks such as model selection and prompt engineering, especially when performance metrics alone cannot distinguish between models. The work is situated in the context of increasing model proliferation and the need for fine-grained comparison beyond aggregate scores.

Innovation

The authors conducted case studies comparing 10 LLMs across 22 Kaggle ML tasks. The results demonstrate that CLIC can effectively distinguish between LLMs based on token-frequency signatures. The decision trees achieved high classification accuracy, indicating that token frequencies alone are sufficient to separate code from different models. The robustness metric revealed that some LLM pairs remain distinguishable even after removing many discriminative tokens, while others become indistinguishable quickly. The concentration metric showed that for some pairs, a few tokens dominate the difference, whereas for others, the difference is spread across many tokens. These findings provide a nuanced view of model similarities and differences. For example, models with similar architectures or training data might exhibit lower robustness or higher concentration, indicating subtle behavioral differences. The visual analytics system enabled the identification of specific tokens that drive differences, such as particular API calls or coding idioms, which can inform prompt engineering strategies.
The evaluation of large language models (LLMs) on coding tasks has primarily focused on performance metrics such as pass@k. As LLMs continue to advance, many models now meet baseline performance requirements, reducing the discriminative power of performance-based evaluation alone. Yet a key question remains largely unexplored: how do LLMs differ in their coding behavior? This paper addresses this gap by proposing CLIC (Code Learning for Identification and Comparison), a visual analytics approach that characterizes LLM coding behavior through token-frequency analysis. The authors argue that understanding behavioral differences is crucial for tasks such as model selection and prompt engineering, especially when performance metrics alone cannot distinguish between models. The work is situated in the context of increasing model proliferation and the need for fine-grained comparison beyond aggregate scores.

CLIC represents each code sample as a feature vector of token frequencies. For a given pair of LLMs, it trains an interpretable decision tree to separate their code sets. The decision tree provides a transparent model of which tokens are most discriminative. Beyond classification accuracy, the authors define two new metrics: robustness and concentration. Robustness measures whether the two LLMs remain distinguishable as their most-discriminative tokens are progressively removed. Formally, let be the set of discriminative tokens sorted by importance. Robustness can be quantified as the area under the curve of classification accuracy as tokens from are removed one by one. Concentration measures whether the difference is driven by a few dominant tokens or spread across many. If is the relative importance of token , concentration can be computed using the Gini coefficient or entropy:

, where is the Shannon entropy. The system supports multi-scale, hypothesis-driven exploration through interactive visualizations, allowing users to navigate the comparison landscape, identify pairs of interest, and drill down into discriminative tokens and their code contexts. The analytical pipeline is illustrated in the following Mermaid diagram:

Why it matters

The paper discusses the implications of token-frequency analysis for understanding LLM coding behavior. The authors argue that performance metrics like pass@k are necessary but not sufficient; behavioral signatures offer complementary insights. The robustness and concentration metrics provide a way to quantify the nature of differences, which can guide model selection: if two models are highly robust in their differences, they may be suited for different tasks, whereas if they are easily confounded, they may be interchangeable. The visual analytics approach supports hypothesis-driven exploration, allowing researchers to form and test hypotheses about model behavior. The authors also note limitations, such as the dependence on tokenization and the potential for token frequencies to be influenced by superficial factors like formatting. Future work could extend CLIC to other modalities or incorporate semantic analysis. Overall, CLIC represents a step towards a more comprehensive evaluation framework for LLMs, emphasizing behavioral characterization alongside performance.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ