WhatsApp is previewing 'Scam Alert,' a new optional feature that uses an on-device machine-learning model to detect scam messages without ever sending message content to Meta — a balancing act between security and the platform's signature end-to-end encryption.
The model is downloaded to the user's device and classifies text locally, with no server involvement. 'No message content leaves the device for classification or is auto-reported to WhatsApp, Meta, or anyone else,' the company says. The system is designed to counter evolving scams, including impersonation and AI-generated lures that trick unsuspecting users.
If the model identifies a likely scam attempt, the user sees a warning inside the chat, invisible to the other party. From there they can block, report or continue the conversation. If a warning is incorrect, users can mark the chat as trusted and future warnings stop; they can also opt in to share the last five messages received to help improve accuracy.
WhatsApp is not rolling the feature out widely yet — it wants to stress-test the system through a limited beta in a few regions first. It is also publishing a technical overview paper and inviting the security community to weigh in. Even the telemetry used to measure false-positive rates is anonymized: aggregate counts are sent to specialized AMD- and Nvidia-powered Trusted Execution Environments that use end-to-end encryption.
The feature lands amid long-running debates over automated scanning in encrypted messengers. Apple abandoned its 2021 plan to scan iPhones for child sexual abuse material after backlash. WhatsApp stresses Scam Alert is optional, can be turned off at any time, and existing user-reporting channels remain unchanged.




