ProPheno: An online dataset for completely characterizing the human protein-phenotype landscape in biomedical literature

2019 
Identifying protein-phenotype relations is of paramount importance for biomedical applications such as uncovering rare and complex diseases. One of the best resources that capture protein-phenotype relationships is the biomedical literature. In this work, we introduce ProPheno 1.0, a comprehensive online dataset composed of human protein/phenotype mentions extracted from the complete corpora of Medline and PubMed Central Open Access. Moreover, it includes co-occurrences of protein-phenotype pairs within different spans of text, such as sentences and paragraphs. We use ProPheno for completely characterizing the human protein-phenotype landscape in biomedical literature. The ProPheno dataset, the reported findings, and the gained insight have implications for (1) biocurators for expediting their curation efforts, (2) researches for quickly finding relevant articles, and (3) text mining tool developers for training their predictive models.
    • Correction
    • Source
    • Cite
    • Save
    • Machine Reading By IdeaReader
    19
    References
    5
    Citations
    NaN
    KQI
    []