Abstract
Retrieval-based augmentation enhances large language models (LLMs) by grounding responses in external knowledge. However, in voice-driven assistants that rely on remote cloud retrieval, open-ended user queries and outsourced processing introduce significant privacy and authorization risks. To address these challenges, we propose SafeRAG, a privacy-preserving retrieval and generation framework that enables secure voice-based interaction over encrypted cloud storage. SafeRAG integrates inner-product functional encryption (IPFE) to support efficient encrypted similarity computation between query and document embeddings without revealing their contents. Each document is protected by an attribute-based access tree that enforces fine-grained authorization, while a Bayesian inference mechanism models the user–assistant dialogue as an adaptive belief-updating process to infer and validate user attributes dynamically. Additionally, a lightweight response sanitization layer applies calibrated noise and semantic abstraction to prevent residual information leakage in generated content. Extensive experiments demonstrate that SafeRAG achieves secure and low-latency retrieval, robust access enforcement, and high-quality voice response generation, offering an effective balance between utility and privacy in remote retrieval-augmented LLMs.
| Original language | English |
|---|---|
| Pages (from-to) | 6211-6224 |
| Number of pages | 14 |
| Journal | IEEE Transactions on Network Science and Engineering |
| Volume | 13 |
| DOIs | |
| State | Published - 2026 |
Keywords
- access control
- Bayesian inference
- privacy-preserving
- Retrieval-augmented generation
Fingerprint
Dive into the research topics of 'SafeRAG: Secure Cloud-Based Retrieval-Augmented Generation for LLM-Empowered Voice Assistants'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver