Bring real-time encrypted voice interaction to Lumo using existing MLS/WebRTC infrastructure
Live Voice Conversation Mode for Lumo — Powered by Proton Meet's Proven Architecture
Summary: Enable bidirectional voice conversation in Lumo's mobile apps — speak through a Bluetooth headset, receive spoken responses via text-to-speech — creating a Gemini Live-style hands-free experience. Proton Meet already solves the hardest engineering challenges for real-time encrypted audio. Extending this to Lumo is a natural evolution, not a fundamental R&D problem.
The Opportunity
Lumo's mobile app supports voice input (speech-to-text) but lacks voice output (text-to-speech) and continuous conversation mode. Users who want hands-free, private AI interaction — while driving, walking, cooking, exercising, or working — have no privacy-preserving alternative to Gemini Live, Siri, or Google Assistant. All of those services harvest voice data for training and profiling.
This is both a feature gap and a market differentiation opportunity: Proton could offer the only AI voice assistant where your spoken conversations are encrypted, never logged, and never used for training.
Technical Feasibility: Proton Meet Already Did the Hard Part
Proton Meet currently delivers real-time, end-to-end encrypted audio for up to 250 participants using a proven, battle-tested architecture:
Component Proton Meet Implementation Applicability to Lumo Voice Mode
Audio capture & encoding WebRTC stack with Opus codec Same stack would power Lumo voice capture/playback
Real-time encryption AES-256-GCM, encrypted on-device before transmission Same encryption protects voice data in transit to/from Lumo backend
Key management MLS (Messaging Layer Security, RFC 9420) — IETF standard co-developed by Proton team members Adapts directly for 1:1 user-to-AI session (simpler than group calls)
Forward secrecy Automatic key rotation per MLS epoch Ensures each voice session is independently secured
Authentication Secure Remote Password (SRP) protocol Same system already used for Proton Account logins and Meet access
Server architecture Selective Forwarding Unit (SFU) via LiveKit Cloud — forwards ciphertext only, never decrypts Lumo backend could process only encrypted requests, same zero-knowledge principle
No-log policy No meeting content, metadata, or participant graphs retained after call ends Same standard applies — voice sessions leave no trace
Infrastructure Distributed European data centers, outside US jurisdiction Same infrastructure hosts Lumo
Bottom line: The cryptographic pipeline for real-time encrypted audio already exists, is audited, is open source, and is in production. Adapting it for a 1:1 user-to-AI voice session is architecturally simpler than what Proton Meet already handles for 250-person group calls with dynamic key rotation.
Proposed Implementation
Phase 1 — Core Voice Mode (6 months):
Tap-to-talk: press to speak, release for response
Streaming TTS output through device speaker or Bluetooth headset
Audio encrypted on-device using MLS-derived keys before sending to Lumo backend
Lumo processes the transcribed text (or encrypted audio) and returns a text response
Response converted to speech via on-device TTS engine
Zero audio or transcription retained after session ends
Phase 2 — Continuous Conversation (12 months):
Hands-free mode: wake word or button activation for back-and-forth dialogue
Conversation state maintained across exchanges within a session
Interruptible responses (user can speak over the AI, like Gemini Live)
Voice activity detection (VAD) to know when the user has finished speaking
Phase 3 — Advanced Features (ongoing):
Voice customization (select TTS voice, speed, language)
Multilingual voice support (leveraging Lumo's existing 11-language support)
Offline TTS via on-device models (no audio leaves the device for output synthesis)
Optional local speech-to-text via on-device Whisper or similar model (input never leaves device)
Voice biometric authentication (optional, user-opt-in, for hands-free session access)
Privacy Architecture: Where Proton Differentiates
Competing voice assistants (Gemini, Siri, Alexa, Google Assistant) share a fundamental flaw: they collect, transcribe, and retain voice data on corporate servers for training, advertising profiling, and analytics. Users have no meaningful control over what happens to their voice after it leaves their device.
Proton's architecture offers a fundamentally different model:
- Local Processing Where Possible
Speech-to-text can run on-device (Android provides native STT APIs; Whisper models run locally on Pixel-class hardware)
Text-to-speech runs locally via Android's native TTS engine or third-party on-device models (e.g., Piper)
When both STT and TTS are local, raw audio never leaves the device — only encrypted text is sent to Lumo's backend for processing
2. Encrypted Transit When Cloud Processing Is Needed
If cloud-based STT/TTS is required (for accuracy, multilingual support, or low-resource devices), audio is encrypted on-device using MLS-derived AES-256-GCM keys before transmission
Lumo's backend processes only encrypted data and discards it immediately — same zero-knowledge architecture as Proton Meet
3. No Retention, No Training
Voice sessions are ephemeral — no audio, transcriptions, or metadata retained after the session ends
Voice data is never used to train models (consistent with Lumo's existing zero-training policy)
No voice fingerprints or biometric profiles stored
4. Transparent and Auditable
All client code is open source (consistent with Lumo's existing open-source mobile apps)
MLS implementation is a public IETF standard — anyone can verify the cryptographic protocol
Independent security audits (Proton is ISO 27001 certified and SOC 2 Type II attested)
Market Positioning
This isn't feature parity with competitors — it's a privacy protection opportunity. Every user who currently uses Gemini Live, Siri, or Alexa because they need hands-free AI is accepting surveillance as the price of convenience. Proton can offer the same convenience with zero surveillance.
Target audiences:
Privacy-conscious consumers who avoid voice assistants because they don't trust Big Tech with their voice data
Professionals handling sensitive information (lawyers, journalists, healthcare workers, executives) who need hands-free AI but can't risk voice data being harvested
Accessibility users who depend on voice interaction and deserve privacy without compromise
Enterprise Lumo for Business customers where voice data retention creates compliance risk under GDPR, HIPAA, and similar regulations
Revenue Potential
Lumo Plus differentiator: Live voice mode as a premium feature driving Plus subscriptions
Lumo for Business upsell: Encrypted voice AI for enterprise — a capability no competitor offers with zero-access encryption
Future hardware partnership opportunity: Licensed secure audio SDK for certified Bluetooth headset manufacturers, bundling with Lumo Plus subscriptions ("Proton Audio Suite")
Community Need
Real-world use cases where private voice AI matters:
Driving: Ask questions, get directions, hear summaries — without sending voice to Google or Apple
Physical work: Cooking, gardening, repairs — hands occupied, need voice interaction
Health and medical: Asking sensitive health questions aloud — voice data is especially intimate
Professional: Dictating notes, drafting documents, brainstorming — work product shouldn't be mined
Accessibility: Visually impaired users or those with motor limitations who rely on voice and deserve privacy equal to sighted users who can type
Recommendation
Prioritize Phase 1 in the current mobile app development cycle. The architecture is proven (Proton Meet), the market demand is demonstrated (Gemini Live adoption), and the privacy differentiation is unmatched. This is a case where Proton can leapfrog competitors by applying existing internal expertise to a new modality.
I am willing to participate in beta testing and provide detailed feedback throughout development.
-
A
commented
This is a critical feature in so many ways; where is the support for this? Please upvote it.