Mining Associations Between Two Categories Using Unstructured Text Data in Cloud
- Yanqing Ji(corresponding author),
- Yun Tian,
- Fangyang Shen,
- John Tran
- ,
- Eastern Washington University,
- New York City College of Technology,
- Frontier Behavioral Health
Abstract
Finding associations between itemsets within two categories (e.g., drugs and adverse effects, genes and diseases) are very important in many domains. However, these association mining tasks often involve computation-intensive algorithms and a large amount of data. This paper investigates how to leverage MapReduce to effectively mine the associations between itemsets within two categories using a large set of unstructured data. While existing MapReduce-based association mining algorithms focus on frequent itemset mining (i.e., finding itemsets whose frequencies are higher than a threshold), we proposed a MapReduce algorithm that could be used to compute all the interestingness measures defined on the basis of a 2 × 2 contingency table. The algorithm was applied to mine the associations between drugs and diseases using 33,959 full-text biomedical articles on the Amazon Elastic MapReduce (EMR) platform. Experiment results indicate that the proposed algorithm exhibits linear scalability.
Bibliographic Information
Output type
Original language
EnglishPages from-to (Number of pages)
Pages 545-550 (6 pages)Publication milestones
- Published - 2018
Publication status
Publisher
Springer VerlagPublication series
- Publication series name: Advances in Intelligent Systems and Computing
ISSN (Print): 2194-5357
Volume: 738
ISBN (Print)
9783319770277Publication IDs
- Scopus: 85045842095
Host publication title
Information Technology - New Generations - 15th International Conference on Information TechnologyHost publication editors
- Shahram Latifi
