Real-Time Speaker Sensitive Audio Gating via Supervised Frame-Level Speaker Recognition
The study presents a modification to the foundation of a noise gate plug-in to create a more intelligent, context-aware system focused on target-speaker speech identification. A supervised, text-independent, and frame-wise personal voice activity detection model based on Ding et al. (2019) is proposed and refined to process audio in real-time with minimal CPU utilization.
The model achieved 87.6% accuracy during batch testing and 72% accuracy on a 5-minute real-world sample under challenging conditions, outperforming a standard noise gate on the same input. Overall, this work demonstrates the feasibility of applying deep learning models in real-time audio systems and demonstrates the potential of personal voice activity detection as a control mechanism for adaptive audio effects, contributing insights to the interdisciplinary fields of computer science and live audio broadcasting.
Technologies: Python, PyTorch, Torch, Google Colab, C++, JUCE Framework
Conference Acceptance: 2026 Joint International Conference on Digital Arts, Media and Technology with ECTI Northern Section Conference on Electrical, Electronics, Computer and Telecommunication Engineering (ECTI DAMT & NCON)