Leverage Temporal Convolutional Network for the Representation Learning of URLs

2019 
Cyber crimes including computer virus/malwares, spam, illegal sales, and phishing websites are proliferated aggressively via the disguised Uniform Resource Locators (URL). Although numerous studies were conducted for the URL classification task, the traditional URL classification solutions retreated due to the hand-crafted feature engineering and the boom of newly generated URLs. In this paper, we study the representation learning of URLs, and explore the URL classification using deep learning. Specifically, we propose URL2vec to extract both the structural and lexical features of URLs, and apply temporal convolutional network (TCN) for the URL classification task. The experimental results show that URL2vec outperforms both word2vec and character-level embedding for URL representation, and TCN achieves the best performance than baselines with the precision up to 95.97%.
    • Correction
    • Source
    • Cite
    • Save
    • Machine Reading By IdeaReader
    20
    References
    1
    Citations
    NaN
    KQI
    []