Computer Science editorial
Testing Deep Learning Library APIs via Cross-Framework Differential Fuzzing
The core problem
Deep learning libraries underpin many safety- and reliability-critical applications, yet existing API-level testing techniques often rely on intra-library properties or CPU--GPU differential oracles and may miss defects that behave consistently across hardware backends. This limitation motivates a cross-framework testing strategy: if two libraries expose APIs intended to implement the same mathematical operation, then discrepancies between their outputs on the same inputs can reveal defects that single-library or hardware-differential testing would overlook.
The paper presents **Xamt**, a cross-framework differential fuzzing approach for deep learning library APIs. Xamt constructs and tests execution-validated groups of APIs intended to implement equivalent operations across seven libraries. It uses explicit API aliases and parameter-role normalization to construct candidate correspondences and validates them through pairwise execution and a group-level behavioral check on canonical ordinary inputs. The resulting groups are explored using variance-guided differential fuzzing with ordinary, boundary, and non-finite inputs. Crash and inconsistency oracles flag executions exhibiting
Innovation
Across the seven libraries, Xamt constructs **676 execution-validated groups** containing **2,563 matched APIs**. Among these, Xamt identifies **72 independently reproduced discrepancy cases**, including **4 crash cases** and **68 output inconsistencies**. Among the **72 developer reports**, **25 have been confirmed**, including **23 that have been fixed**.
The results demonstrate that cross-framework differential fuzzing can uncover defects that are consistent across hardware backends and therefore missed by CPU--GPU differential oracles. The high confirmation rate (25 of 72 reports confirmed, with 23 fixes) indicates that the discrepancies are actionable and relevant to library maintainers.
Key quantitative outcomes:
- Execution-validated groups: 676
- Matched APIs: 2,563
- Reproduced discrepancy cases: 72
- Crash cases: 4
- Output inconsistencies: 68
- Developer reports: 72
- Confirmed reports: 25
- Fixed reports: 23
Why it matters
The Xamt results highlight the value of cross-framework differential testing for deep learning library APIs. By constructing execution-validated groups of APIs intended to implement equivalent operations, Xamt can detect inconsistencies that are not observable through intra-library properties or CPU--GPU differential oracles. The use of explicit API aliases and parameter-role normalization enables correspondence construction across libraries with different naming and parameter conventions, while execution validation and group-level behavioral checks reduce false positives.
Variance-guided differential fuzzing with ordinary, boundary, and non-finite inputs proves effective: the 72 reproduced discrepancy cases include both crashes and output inconsistencies, and the 25 confirmed developer reports (23 fixed) suggest that the findings are meaningful to developers. The approach is complementary to existing testing techniques and can be integrated into library CI pipelines to catch cross-framework inconsistencies early.
Limitations include reliance on documented semantics and aliases for correspondence construction, and the tolerance used in validation may need tuning per operation. Future work could extend Xamt to more libraries, incorporate automated alias discovery, and explore adaptive input generation strategies.
In summary, Xamt demonstrates that cross-framework differential fuzzing is a practical and effective method for testing deep learning library APIs, yielding reproducible discrepancies and actionable developer reports.
Who should read this
Opening member contentโฆ