Extend Lumo to have integrated TTS/STT for full Voice Interaction
For general Chat, its really nice to be able to have a conversation using voice rather than keyboard. And while I know that the Android/IOS apps offers STT, its unfortunate that I have to use the phones TTS to complete the loop. For my tablet and/or webbrowser a full TTS/STT feature would be nice.
-
Matthew Kowalski commented
If Lumo ever gets the ability to read back TTS/STT, it might be a fun feature to give the voice a somewhat feline feel (maybe the occasional mreow/purr at the beginning of the response or a <tail twitch> in the text response). It'd make for a warmer personality interface.
-
Jazzy
commented
Privacy-first voice integration — Proton's unique advantage
A native Lumo TTS/STT would solve two problems at once:
1. Privacy: Currently, voice input on mobile relies on Apple/Google STT, which raises legitimate concerns about what happens to voice data — dialect, speech patterns, intonation. A Proton-native, zero-access-encrypted voice system would keep that data where it belongs: with the user.
2. Platform independence: A built-in system wouldn't just benefit iOS/Android users — it would bring full voice interaction to de-Googled phones, Linux tablets, and the web app, where no native OS STT/TTS is available. This aligns perfectly with Proton's mission of privacy for everyone, not just those locked into mainstream ecosystems.
Additionally, the points raised by others here are spot on: hands-free conversation mode (automatic end-of-speech detection, continuous dialogue toggle), and curated spoken responses (summarizing tables rather than reading them cell by cell) would make voice a first-class way to interact with Lumo — not just an accessibility afterthought.
ChatGPT and Claude already have full voice mode. Proton has the opportunity to do it better — private, encrypted, and platform-independent.
Critical for accessibility and inclusion
Beyond privacy and platform independence, this feature is critical for accessibility. Without integrated TTS/STT, Lumo is effectively unusable for blind or visually impaired users who rely on screen readers and voice interaction to navigate digital spaces. True inclusivity means ensuring that privacy-focused tools are accessible to everyone, regardless of ability. Making voice a first-class citizen isn't just a convenience—it's a necessity for equal access.
-
Francoise
commented
Absolutely. The ability to have AI read back answers is an essential function and a dealbreaker to really commit to Lumo. Please add this ASAP.
-
Marc
commented
I couldn't agree more. It would really make life easier to continue working hands free, continuously improving production.
-
Proton Nutzer
commented
Even for the use on iPhones the local STT does not cut it,
becase you loose alle the information that is carried by your intonation (you even loose the interpunctuation).
Also if the STT is done locally, Lumo cannot adjust to your dialect and speech patters (the things we're trying to keep out of the hands of the data brokers)..
On the reverse path, having to use the local TTS ist just a bad joke.
I want to have a proper conversation with Lumo, without having to press a button to read the answers.
So instead of just the "microphone STT button" and a "read replies TTS button", an additional selector should be there,
something like a "conversation toggle switch", that activates both, until you turn it off..
And i.e. when replying by voice,
Lumo should not just read what's on the screen, cuz. if lumo starts reading a table, that does not carry very well in speech.
So Lumo should point you to the full table in the chat, but curate what it replies for the auditive communication channel. -
Jolene Cook commented
Ditto!
-
Tony
commented
Not only that, but the STT doesn't realise that we finished our question/sentence and waits till we manually click send, making the experience not hands-free even with the use of the OS TTS. I think this is the future and there needs to be full voice mode like Grok, ChatGPT..etc
-
Ananda
commented
I think there needs to be a full Voice Mode (Like Claude and ChatGPT) and also an icon below every response from Lumo that enables reading each message individually! Those two options together form a complete voice integration which increases usability and accessibility! Which is a win-win situation :) Agreed with John Doe, such feature is absolutely necessary, otherwise it's like having a gigantic mansion with only one entrance at the back ;) I use AI assistants on voice mode 95% of the time! Thanks!
-
John Doe
commented
Please - This is severely needed. An LLM without voice mode is only half useful. Ideally, it would be a full conversation mode where you are able to speak, the AI realize you are done speaking and reply, and then you are able to reply...etc. All without clicking the microphone icon multiple times or clicking "play" on the AI's response.