Trie-based rule processing for clinical NLP: A use-case study of n-trie, making the ConText algorithm more efficient and scalable

Jianlin Shi; John F Hurdle

doi:10.1016/j.jbi.2018.08.002

Trie-based rule processing for clinical NLP: A use-case study of n-trie, making the ConText algorithm more efficient and scalable

J Biomed Inform. 2018 Sep:85:106-113. doi: 10.1016/j.jbi.2018.08.002. Epub 2018 Aug 6.

Authors

Jianlin Shi¹, John F Hurdle²

Affiliations

¹ Department of in Biomedical Informatics, University of Utah, Salt Lake City, UT, USA. Electronic address: jianlin.shi@utah.edu.
² Department of in Biomedical Informatics, University of Utah, Salt Lake City, UT, USA. Electronic address: john.hurdle@utah.edu.

Abstract

Objective: To develop and evaluate an efficient Trie structure for large-scale, rule-based clinical natural language processing (NLP), which we call n-trie.

Background: Despite the popularity of machine learning techniques in natural language processing, rule-based systems boast important advantages: distinctive transparency, ease of incorporating external knowledge, and less demanding annotation requirements. However, processing efficiency remains a major obstacle for adopting standard rule-base NLP solutions in big data analyses.

Methods: We developed n-trie to specifically address the token-based nature of context detection, an important facet of clinical NLP that is known to slow down NLP pipelines. N-trie, a new rule processing engine using a revised Trie structure, allows fast execution of lexicon-based NLP rules. To determine its applicability and evaluate its performance, we applied the n-trie engine in an implementation (called FastContext) of the ConText algorithm and compared its processing speed and accuracy with JavaConText and GeneralConText, two widely used Java ConText implementations, as well as with a standalone machine learning NegEx implementation, NegScope.

Results: The n-trie engine ran two orders of magnitude faster and was far less sensitive to rule set size than the comparison implementations, and it proved faster than the best machine learning negation detector. Additionally, the engine consistently gained accuracy improvement as the rule set increased (the desired outcome of adding new rules), while the other implementations did not.

Conclusions: The n-trie engine is an efficient, scalable engine to support NLP rule processing and shows the potential for application in other NLP tasks beyond context detection.

Keywords: Algorithms; Data accuracy; Medical informatics applications; Natural language processing.

Publication types

Comparative Study
Evaluation Study
Research Support, N.I.H., Extramural
Research Support, Non-U.S. Gov't

MeSH terms

Algorithms*
Computational Biology
Databases, Factual
Humans
Machine Learning
Natural Language Processing*

Grants and funding

R01 LM010981/LM/NLM NIH HHS/United States