Jadwal Sholat

Memuat jadwal sholat…

Computer Science editorial

Open AccessOA2026

Custom Named Entity Recognition and Topic Classification for Global Health Publications

An empirical study on adapting NLP models for global health literature under resource constraints
Genis Skura; Antoine Geissbühler; Jean-Luc Falcone· 2026· DOI 10.48550/arXiv.2609.24625

The core problem

The increasing volume of global health literature demands automated tools for indexing and knowledge extraction. However, many settings face limited annotated data and computational resources. This thesis addresses the challenge of selecting and adapting natural language processing (NLP) models for global health publications under such constraints. It focuses on three tasks: semantic tag discovery, named entity recognition (NER), and multi-label topic classification. The goal is to provide an empirical basis for building knowledge systems that balance domain specialization, accuracy, and computational efficiency.

Innovation

Key findings from the experiments:

- **Semantic Tag Discovery**: Broader vocabulary coverage does not necessarily yield more useful domain-specific associations. Qualitative evaluation indicates that specialized corpora of moderate size may suffice for meaningful tag discovery.

- **Named Entity Recognition**: Under lenient scoring, the RoBERTa transformer achieves 0.80 micro-, outperforming convolutional models (0.65-0.69). However, its inference time (82 seconds) is over 13 times longer than convolutional models (5-6 seconds). The integrated disease recognizer achieves 81.33% test on the NCBI Disease Corpus.

- **Topic Classification**: BART-MNLI zero-shot inference achieves 95.2% single-label accuracy, significantly higher than MiniLM few-shot (59%). For multi-label classification, BART-MNLI reaches 88% accuracy versus 32% for MiniLM under partly manual assessment. However, BART-MNLI's higher inference cost limits practical integration.

These results highlight trade-offs between accuracy and computational efficiency.

The increasing volume of global health literature demands automated tools for indexing and knowledge extraction. However, many settings face limited annotated data and computational resources. This thesis addresses the challenge of selecting and adapting natural language processing (NLP) models for global health publications under such constraints. It focuses on three tasks: semantic tag discovery, named entity recognition (NER), and multi-label topic classification. The goal is to provide an empirical basis for building knowledge systems that balance domain specialization, accuracy, and computational efficiency.
The research comprises three experimental strands:

Why it matters

The thesis demonstrates that domain specialization and lightweight adaptation offer practical value in resource-constrained environments. For NER, fine-tuned convolutional models provide a balance between accuracy and speed, while transformer models justify their higher inference costs when maximum accuracy is critical. The disease recognizer's strong performance (81.33% ) shows that targeted fine-tuning can yield robust results even with limited data. For topic classification, zero-shot BART-MNLI excels in accuracy but may be impractical for large-scale or real-time applications due to inference cost. Few-shot MiniLM, while less accurate, is more efficient. The findings suggest a hybrid approach: use lightweight models for high-throughput tasks and reserve transformers for tasks where accuracy is paramount. The pipeline's integration of PDF preprocessing, entity filtering, and MeSH enrichment supports document-level indexing, providing a foundation for knowledge systems in global health. Future work could explore model distillation and further domain adaptation to reduce inference costs while maintaining accuracy.

Who should read this

CS practitioners and researchers

Opening member content…