Computer Science editorial
Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models
The core problem
Innovation
CLIC represents each code sample as a feature vector of token frequencies. For a given pair of LLMs, it trains an interpretable decision tree to separate their code sets. The decision tree provides a transparent model of which tokens are most discriminative. Beyond classification accuracy, the authors define two new metrics: robustness and concentration. Robustness measures whether the two LLMs remain distinguishable as their most-discriminative tokens are progressively removed. Formally, let be the set of discriminative tokens sorted by importance. Robustness can be quantified as the area under the curve of classification accuracy as tokens from are removed one by one. Concentration measures whether the difference is driven by a few dominant tokens or spread across many. If is the relative importance of token , concentration can be computed using the Gini coefficient or entropy:
Why it matters
Who should read this
Opening member contentโฆ