Transparency Gap in Lumo — Model Disclosure Consistency Issue
Dear Proton Team,
I'm writing as a supporter of Proton's privacy mission who deeply appreciates what you've built with Lumo. The zero-access encryption architecture, European datacenter infrastructure, and commitment to keeping conversations off training pipelines are exactly what the AI privacy space needs. However, I've identified a significant inconsistency in transparency that undermines this otherwise exceptional privacy-first positioning.
THE CORE ISSUE
Currently, Proton markets Lumo as using "open-source language models" without disclosing which specific base models or model families power the service. This creates a transparency gap between your stated values and your actual disclosure practices.
WHAT IS CURRENTLY DISCLOSED:
- "Open-source language models" (vague category)
- European datacenters
- Zero-access encryption
- No logs, no training on user data
WHAT IS NOT DISCLOSED:
- Base model name/family (e.g., Llama 3.2, Mistral Large, etc.)
- License type (Apache 2.0, Llama Community License, etc.)
- Fine-tuning methodology description
- Version/commit references for reproducibility
INDUSTRY BENCHMARK COMPARISON
Company: DuckDuckGo (Duck.ai)
Disclosure: OpenAI GPT-5.4, Anthropic Claude 4.5 Haiku, Meta Llama 3.3, Mistral Small 4 — all publicly named
Company: Mullvad VPN
Disclosure: Full app source code on GitHub (GPLv3), WireGuard/OpenVPN protocols, annual transparency reports, published police raid documentation
Company: Quad9 DNS
Disclosure: All 15+ threat-intelligence partners listed by name, quarterly transparency reports, infrastructure maps, human rights impact statements
Company: Proton Lumo
Disclosure: "Open-source models" — no specific names, licenses, or technical lineage
The question is straightforward: If transparency is a core differentiator, why does it stop at the encryption layer?
WHY THIS MATTERS FOR TRUST
Security Researchers: Can they audit the actual ML pipeline for vulnerabilities or data leakage risks?
Enterprise Buyers: What license governs derivative use of outputs? Are there compliance implications?
Technical Users: Is this fine-tuned Llama vs. Mistral vs. Gemma? These have vastly different capabilities.
Privacy Advocates: Does the model's original training data introduce indirect data exposure risks?
Community Contributors: How can developers build extensions or integrations without knowing the foundation?
Being "open about being open-source" and being actually specific are two different things. The latter builds verifiable trust; the former relies on blind faith.
LESSONS FROM PRIVACY-FOCUSED PEERS
Mullvad's Approach:
"All of Mullvad's apps are open source. This means anyone can inspect the code, verify what it does, and ensure there are no backdoors or vulnerabilities."
They don't just claim security—they make it auditable.
Quad9's Approach:
"Quad9 integrates threat intelligence from multiple industry leaders... We publish quarterly transparency reports showing zero data requests received since 2017."
They don't just claim privacy—they document proof points.
Proton's Opportunity:
You could match this standard with a simple disclosure like:
"Lumo's core capabilities derive from fine-tuned versions of [Model Family A] and [Model Family B], with proprietary privacy-enhancing modifications."
This doesn't reveal trade secrets or compromise security—it just satisfies legitimate user expectations aligned with your transparency brand promise.
SPECIFIC RECOMMENDATIONS
- Publish a blog post listing base model names (Low effort, High trust impact)
- Add a model card to Lumo documentation (Low effort, High trust impact)
- Publish fine-tuning methodology overview (Medium effort, High trust impact)
- Include model lineage in security model docs (Medium effort, High trust impact)
- Release evaluation benchmarks for comparison (Medium effort, Very High trust impact)
Even starting with just naming the base model families would address 80% of the transparency gap while preserving operational flexibility.
ADDRESSING POTENTIAL CONCERNS
Concern: Competitors will copy our fine-tuning
Rebuttal: Naming base models doesn't reveal fine-tuning parameters or training data
Concern: We'll lock ourselves into specific versions
Rebuttal: You can disclose "Llama 3.2 family" or "Mistral Large class" without binding to exact versions
Concern: Security through obscurity
Rebuttal: Security-by-design is more trustworthy than security-by-secret. Auditable systems build trust.
Concern: Marketing differentiation
Rebuttal: True capability differentiation comes from performance + privacy, not model obscurity.
None of these justify zero disclosure about base models when transparency is a core marketing pillar.
BOTTOM LINE REQUEST
If Proton truly believes in transparency as a foundational value (not just a marketing tagline), please consider:
- Disclosing base model families used in Lumo
- Including licensing information for those base models
- Publishing a high-level architecture diagram showing how models integrate with encryption
- Adding transparency reports for model updates and incidents
This would align Lumo's disclosure standards with:
- Your own encryption/security documentation (which is excellent)
- Peer privacy companies like Mullvad and Quad9
- User expectations for a "transparent privacy-first AI"
THANK YOU
I understand product teams balance competing priorities—security, competitiveness, development velocity, and transparency. But for companies whose core brand promise is privacy and transparency, consistency across all dimensions matters. Every undisclosed detail erodes the cumulative trust earned by every disclosed promise.
Thank you for reading this feedback and for building tools that give users real control over their digital lives. I look forward to seeing Lumo continue to set industry standards.
Best regards,
Security Researcher