Skip to main navigation Skip to search Skip to main content

Fine-Grained Detection and Analysis of Unknown Encrypted Malicious Traffic From Mixed Noisy Labels

  • Qianwei Meng
  • , Qingjun Yuan*
  • , Meng Shen
  • , Siqi Lu
  • , Guangsong Li
  • , Jing Tao
  • , Yong Yu
  • , Yongjuan Wang
  • *Corresponding author for this work
  • Information Engineering University
  • Ministry of Education in China
  • Beijing Institute of Technology
  • Xi'an Jiaotong University
  • Shaanxi Normal University

Research output: Contribution to journalArticlepeer-review

Abstract

The efficacy of deep learning-based Network Intrusion Detection Systems (NIDS) is critically constrained by the availability of high-quality labeled data. In real-world environments, datasets often suffer from mixed label noise—consisting of both closed-set and open-set noise—which significantly distorts decision boundaries and leads to critical security misses for previously unknown attacks. Achieving robust classification of known traffic while accurately detecting unknown threats in the presence of such mixed noise remains a key challenge. In this paper, we introduce Sieve, a robust framework designed for the fine-grained detection and analysis of unknown encrypted malicious traffic in mixed noise conditions. Sieve consists of three collaborative modules: (i) a noise-resilient label correction module that filters out mixed noise using neighbor consistency metrics and confidence-based subset expansion; (ii) a post-hoc detection module that employs Mahalanobis distance in a purified, compact feature space to identify unknown traffic; and (iii) an unknown traffic labeling module that utilizes semi-supervised clustering to facilitate efficient updates of the dataset. Empirical evaluations across four public datasets demonstrate that Sieve significantly outperforms state-of-the-art methods in both known-class classification and unknown-threat detection. Notably, on the Mal_TLS2023 dataset under 50% noise conditions, Sieve maintains robust performance, with accuracy and F1 scores exceeding 94%. Further theoretical analyses corroborate the superiority of the Sieve framework.

Original languageEnglish
JournalIEEE Transactions on Dependable and Secure Computing
DOIs
Publication statusAccepted/In press - 2026
Externally publishedYes

Keywords

  • Deep learning
  • malicious traffic detection
  • noisy label
  • unknown attack

Fingerprint

Dive into the research topics of 'Fine-Grained Detection and Analysis of Unknown Encrypted Malicious Traffic From Mixed Noisy Labels'. Together they form a unique fingerprint.

Cite this