Computer Science editorial
Open AccessOA2026
Harnessing LLMs for Document-Guided Fuzzing of Python Libraries
VistaFuzz: Extracting parameter specifications from API documentation to generate constraint-satisfying inputs and uncover real-world bugs
Bin Duan; Tarek Mahmud; Meiru Che; Yan Yan; Naipeng Dong; Dan Dongseong Kim; Guowei Yangยท 2026ยท DOI 10.48550/arXiv.2608.11744
The core problem
Python libraries are foundational to deep learning, scientific computing, data analysis, and computer vision, so their reliability directly affects downstream applications. Testing their APIs is challenging because valid inputs must satisfy both per-parameter constraints (e.g., types, ranges, formats) and dependencies among parameters (e.g., one argument must be set if another is absent). Existing fuzzing approaches either leave such constraints implicit in generated programs or rely on library-specific parsing rules, limiting generality and effectiveness. This paper introduces **VistaFuzz**, a document-guided fuzzing technique that leverages a locally served open-sourced LLM to extract parameter specifications from API documents and generate inputs that satisfy both parameter constraints and inter-parameter dependencies. The work targets the gap between generic program generation and the need for semantically valid API calls, aiming to improve valid input generation and bug discovery across diverse Python libraries.
Innovation
The authors evaluate VistaFuzz on **7,718 APIs across twelve Python libraries**. They find that inter-parameter relationships occur in **40.1%** of tested APIs. When dependency resolution is disabled, the valid generation rate on those APIs drops from **above 95%** to **31.6%โ52.8%**, demonstrating the importance of resolving inter-parameter dependencies. VistaFuzz reports **74 issues**, of which **43 have been confirmed by developers** and **29 have been fixed**. These results indicate that document-guided, constraint-aware fuzzing can generate highly valid inputs and uncover real bugs in widely used libraries. The substantial drop in valid generation when dependencies are ignored highlights that per-parameter constraints alone are insufficient for effective API testing.
Python libraries are foundational to deep learning, scientific computing, data analysis, and computer vision, so their reliability directly affects downstream applications. Testing their APIs is challenging because valid inputs must satisfy both per-parameter constraints (e.g., types, ranges, formats) and dependencies among parameters (e.g., one argument must be set if another is absent). Existing fuzzing approaches either leave such constraints implicit in generated programs or rely on library-specific parsing rules, limiting generality and effectiveness. This paper introduces **VistaFuzz**, a document-guided fuzzing technique that leverages a locally served open-sourced LLM to extract parameter specifications from API documents and generate inputs that satisfy both parameter constraints and inter-parameter dependencies. The work targets the gap between generic program generation and the need for semantically valid API calls, aiming to improve valid input generation and bug discovery across diverse Python libraries.
VistaFuzz operates in three main stages: (1) **Documentation parsing and specification extraction**, (2) **Constraint-aware input generation**, and (3) **Fuzzing and issue reporting**. Given an API, VistaFuzz retrieves its documentation and prompts a locally served open-source LLM to extract a structured specification of parameters, including types, allowed values, and inter-parameter dependencies. These specifications are then used to guide the generation of valid inputs. The LLM is constrained to produce inputs that satisfy the extracted constraints, and a dependency resolver ensures that inter-parameter relationships are respected. The generated inputs are executed against the library under test, and any failures are triaged into issue reports. The approach is library-agnostic, requiring no library-specific parsing rules, and runs entirely locally to avoid reliance on external services.
Why it matters
The findings underscore that inter-parameter dependencies are pervasive in Python library APIs and that ignoring them severely degrades input validity. VistaFuzz's use of a locally served open-source LLM avoids dependence on proprietary services and library-specific parsing rules, making the approach broadly applicable. The high valid generation rate (above 95% with dependency resolution) suggests that LLM-based specification extraction can effectively capture complex constraints from documentation. The 74 reported issues, with 43 confirmed and 29 fixed, provide evidence of real-world impact. However, the study also implies that documentation quality and LLM extraction accuracy are critical; incomplete or ambiguous docs could lead to missed constraints. Future work may extend VistaFuzz to other languages and ecosystems, and explore combining static analysis with LLM-based extraction to further improve coverage and precision.
Who should read this
CS practitioners and researchers
Opening member contentโฆ