Voice AI that never leaves the device.
On-device voice intelligence — transcription, synthesis, wake-word, and more. From wake-word detection to speech synthesis, every engine runs on the device with zero cloud dependency. Arabic, Hindi and Urdu are built into the foundation alongside English.
Built in Rust by Sajdak Group Holdings W.L.L. Hundreds of tests passing. Patent portfolio filed.
- Hey SajEnglish
- يا ساجArabic
- हे साजHindi
- ہے ساجUrdu
- Wake-word inference
- <1 ms
- Wake-word model size
- <1 MB
- Trained models
- 70+
- Languages: English, Arabic, Hindi, Urdu
- 4
Why Saj Speak
Voice AI that respects your users.
Three principles, no compromise.
-
Private by design
Your voice never leaves your device. No cloud processing, no data retention, no telemetry. GDPR, HIPAA, and NESA compliant by architecture — not by policy.
-
Arabic-first
Not a translation. Not an afterthought. Arabic, Hindi, and Urdu are built into the foundation alongside English.
-
Every platform
iOS, Android, Linux, Windows, macOS, Raspberry Pi, WASM, MCU. A Rust core with native bindings — deploy to phones, speakers, embedded devices, and browsers.
Engines
Every engine runs on the device.
No cloud round-trip. Ship voice features that work offline, everywhere. Each engine is its own Rust crate.
Listen
Wake word saj-wake
Custom keyword detection. Lightweight neural architecture, under 1 MB models. English, Arabic, Hindi, Urdu.
Voice activity saj-detect
Neural voice activity detection with sub-millisecond latency.
Streaming speech-to-text saj-listen
Real-time transcription. English, Arabic, Hindi, Urdu.
Batch speech-to-text saj-scribe
File transcription with a split encoder and decoder.
Speech intent saj-intent
Direct intent extraction. Voice commands to actions, skipping the text. English, Arabic.
Speak
Text-to-speech saj-speak-tts
Natural on-device synthesis. English, Arabic, Hindi.
Voice cloning saj-clone
Clone a voice from samples, on-device. English, Arabic.
Who is speaking
Speaker ID saj-voice
Speaker embeddings for verification and identification.
Diarization saj-who
Real-time speaker segmentation. Who spoke when.
Noise suppression saj-clean
Neural noise removal. Lightweight.
Connect
Pipeline saj-pipeline
Unified voice interaction: wake, voice activity, speech-to-text, intent, text-to-speech.
Neural codec saj-codec
Neural speech codec with quality profiles from 1.2 to 6.0 kbps. Real-time on CPU.
Multimodal sensing saj-sense
Unified token streams with per-modality encryption.
Encrypted comms saj-link
End-to-end encrypted messaging. Sealed sender.
Self-hosted server saj-server
REST and WebSocket, with an OpenAI- and Deepgram-compatible API. Multi-tenant. Docker.
Native Arabic
Arabic in the foundation, dialect by dialect.
420M+ Arabic speakers. 1.09 billion Arabic, Hindi, and Urdu speakers combined.
| Dialect | Coverage | Status |
|---|---|---|
| الفصحى Modern Standard |
Wake word, STT, TTS, intent. The formal register of 420M+ speakers. | Supported |
| الخليجي Gulf |
Wake word, STT, TTS. Tuned for UAE, Saudi, Bahrain, Kuwait, Qatar, Oman. | Supported |
| المصري Egyptian |
100M+ speakers. The most widely understood Arabic dialect worldwide. | Planned |
| الشامي Levantine |
Syria, Lebanon, Jordan, Palestine. 30M+ speakers. | Planned |
Arabic wake words available
Enterprise MENA voice AI
Banks, telcos, and government agencies across the Gulf need voice AI that processes Arabic on-device — with dialect awareness, sovereign data processing, and NESA compliance built in.
Access
How to get Saj Speak.
Saj Link is the application you use every day. Saj Speak, Saj See, and Saj Codec are internal capability modules that surface through Saj Sense Business+, the hosted developer API, and OEM/Embedded engagements — never sold as separate brands.
-
Saj Sense Business+
Saj Speak surfaces in the Saj Sense Business+ tier.
-
Hosted developer API
Saj Speak is also available through the hosted developer API.
-
OEM / Embedded
Saj Speak inside your own device or product, through an OEM/Embedded engagement.
On-device
Why on-device?
Cloud voice APIs charge per minute, leak data, and add latency. On-device fixes all three.
-
Total privacy
Zero audio data leaves the device. No cloud processing. No data retention.
-
Sub-millisecond wake word
Sub-millisecond wake-word inference against a 200–500 ms cloud round-trip. Real-time voice interaction without waiting for the network.
-
No per-minute charges
No audio-hour billing. Your models run on your hardware.
Built on Saj Speak
Get Saj Link.
Encrypted team communication with Bella AI built in. Post-quantum protocol built, rolling out. Available now for macOS and web.
-
End-to-end encryption
Sealed sender. End-to-end encryption for messages and shared media. Call audio is relayed, not yet end-to-end encrypted.
-
Voice intelligence
On-device transcription, translation, and search. Dialect-aware.
-
Arabic-first design
Native RTL, dialect awareness, voice messages as rich cards.