[Online]Context-Aware Real-Time Audio and Image Toxicity Moderation via Multi-Agent Reinforcement Learning for Cybersecurity in Networked Communication Platforms

Context-Aware Real-Time Audio and Image Toxicity Moderation via Multi-Agent Reinforcement Learning for Cybersecurity in Networked Communication Platforms
ID:67 Submission ID:442 View Protection:ATTENDEE Updated Time:2026-07-25 18:02:18 Hits:33 Online

Start Time:2026-07-30 15:25 (Asia/Kolkata)

Duration:15min

Session:[S3] Cyber Security » [S3-2] Cyber Security

Video No Permission Presentation File Attachment File

Tips: The file permissions under this presentation are only for participants. You have not logged in yet and cannot view it temporarily.

Abstract

Toxic content propagation across networked communication

platforms constitutes an emerging cybersecurity

challenge, as real-time audio and image channels introduce

attack surfaces that bypass conventional text-only defences. Automated

content moderation in online communication platforms

increasingly requires coverage beyond text, as toxic content is

frequently delivered through audio and image channels that textonly

systems cannot address. This paper presents a dual-modality

toxic content detection and moderation pipeline operating across

audio and image inputs in real time. For audio, the Roblox

voice-safety-classifier-v2, a WavLM transformer pretrained on

over 100,000 hours of real gaming voice chat, generates sixdimensional

toxicity probability scores mapped to the Jigsaw

taxonomy via a novel cross-taxonomy semantic alignment, with

faster-whisper providing word-level timestamps for surgical muting

of precisely the toxic speech segments rather than entire

clips. For image moderation, CLIP ViT-B/32 encodes images

against toxicity-describing natural language prompts to produce

a nine-dimensional feature vector, with flagged content reposted

as blurred spoilers. Proximal Policy Optimisation reinforcement

learning agents trained with an asymmetric reward structure

penalising false negatives over false positives achieve 97.5%

accuracy on unseen audio evaluation data with 91.0% reward efficiency,

and 87.1% reward efficiency on image data, significantly

outperforming rule-based, random, always-mute, and alwaysallow

baseline policies. A per-user cross-modal trust score system

with progressive escalation is deployed as a real-time automated

Discord moderation bot validated through live user interactions.

Keywords
audio toxicity detection, image moderation, reinforcement learning, proximal policy optimisation, WavLM, CLIP, gaming platforms, content moderation, cybersecurity
Speaker
Vemula Yashodha
Student Amrita University

Submission Author
Vemula Yashodha Amrita University
Sri Ramya Divakarla Amrita University
Vudumala Rupa Manogna Amrita University
SUSMITHA VEKKOT Amrita School of Engineering
Comment submit
Verification code Change another
All comments
Log in Sign up Registration Submission