Computer Science editorial
Custom Named Entity Recognition and Topic Classification for Global Health Publications
The core problem
Innovation
Key findings from the experiments:
- **Semantic Tag Discovery**: Broader vocabulary coverage does not necessarily yield more useful domain-specific associations. Qualitative evaluation indicates that specialized corpora of moderate size may suffice for meaningful tag discovery.
- **Named Entity Recognition**: Under lenient scoring, the RoBERTa transformer achieves 0.80 micro-, outperforming convolutional models (0.65-0.69). However, its inference time (82 seconds) is over 13 times longer than convolutional models (5-6 seconds). The integrated disease recognizer achieves 81.33% test on the NCBI Disease Corpus.
- **Topic Classification**: BART-MNLI zero-shot inference achieves 95.2% single-label accuracy, significantly higher than MiniLM few-shot (59%). For multi-label classification, BART-MNLI reaches 88% accuracy versus 32% for MiniLM under partly manual assessment. However, BART-MNLI's higher inference cost limits practical integration.
These results highlight trade-offs between accuracy and computational efficiency.
Why it matters
Who should read this
Opening member content…