23 open models · WebGPU accelerated · No ad tracking
Free AI Chat, Images & Voice — 100% In Your Browser
No signup. Your files never leave your device. Runs locally after setup. Prepare and test before going offline.
- WebGPU accelerated
- WASM fallback
- 23 open models
- Local processing
Choose a tool
Start where you are
Chat
Think, write, and explore ideas with local LLMs.
- Qwen 3.5 & more
- File knowledge & voice mode
Image
Generate, refine, and transform images.
- Text-to-image, depth & upscale
- One-click background removal
Audio
Create speech or turn recordings into text.
- 27 neural voices
- Speech clips & timestamped transcripts
What would you like to do?
What's inside
Three studios, one browser tab
Chat
-
Qwen 3.5 & more
Switch models mid-conversation
-
Thinking mode
Watch compatible models reason before answering
-
Chat with your files
PDF & text knowledge via local RAG
-
Voice in, voice out
Local transcription input, Kokoro read-aloud replies
Image
-
Text-to-image
SD-Turbo fast drafts or SDXL-Lightning quality
-
Image to text
Extract, copy, and ask about printed text
-
Depth maps
Depth Anything 3 scene structure
-
Upscale
Swin2SR 2×/4× super-resolution
-
Remove background
BiRefNet one-click cutouts
Audio
-
Neural TTS
Kokoro-82M with 27 voices & 3 speeds
-
Transcribe
Moonshine English & multilingual Whisper
-
Clip history
Playback, timestamps, and subtitle exports
-
Record & import
Live mic capture with level meter
How it works
Local in three steps
- 01
Pick a tool
Chat, Image, or Audio. No account, no setup, nothing to configure.
- 02
Model downloads once
Open weights come straight from Hugging Face and are cached in your browser.
- 03
Runs locally & privately
Inference happens on your GPU or CPU. AI content stays on your device unless you activate an external provider.
Private by design
Runs on your device. Proof, live.
Models cached
Cached model files — processing runs on your device
Model weights download once from Hugging Face, then inference runs locally. Test cached tools before disconnecting.
- Chrome / Edge WebGPU
- Firefox WASM fallback
- Safari WASM fallback
WebGPU unlocks GPU-accelerated inference. Firefox and Safari can run most tools through the WASM CPU path — image generation (SD-Turbo, SDXL-Lightning) and the largest chat models require WebGPU.
Full model matrix 23 models
| Model | Area | Parameters | Download | VRAM |
|---|---|---|---|---|
| SmolLM2 360M | chat | 360M (q4f16_1) | ~220 MB | 400 MB |
| LFM2.5 350M | chat | 350M (q4f16) | ~210 MB | 450 MB |
| Qwen 3.5 0.8B | chat | 0.8B (q4f16) | 430 MB | ~1.6 GB |
| Qwen 3.5 2B | chat | 2.3B (q4f16) | 1.1 GB | ~2.2 GB |
| Qwen 3.5 4B | chat | 4B (q4f16) | 2.3 GB | 3.9 GB |
| Qwen 3.5 9B | chat | 9B (q4f16) | 4.8 GB | 6.4 GB |
| SD-Turbo | image | Single-step Turbo | 2.3 GB | 1.2 GB |
| SDXL-Lightning | image | 2.6B UNet (INT4) | ~3.6 GB | 6.0 GB |
| Depth Anything 3 Small | image | ~26M | 105 MB | 300 MB |
| Depth Anything 3 Base | image | ~103M | ~413 MB | 800 MB |
| Swin2SR Lightweight 2x | image | 1.01M · fp32 | 8.08 MB | 250 MB |
| Real-ESRGAN General x4v3 | image | 1.21M · fp32 | 4.87 MB | 300 MB |
| Swin2SR Classical 4x | image | 12.2M · q8 | 21.6 MB (q8) | 500 MB |
| Fast | image | BiRefNet Lite (512) | ~99 MB | 300 MB |
| Quality | image | BEN2 Base | ~219 MB | 600 MB |
| PP-OCRv5 Mobile | image | Det 4.7M + Rec 15.8M | ~21.3 MB | 150 MB |
| Kokoro 82M | audio | 82M | 86 MB | 200 MB |
| Moonshine Tiny | audio | 27M | ~53 MB | 150 MB |
| Moonshine Base | audio | 61M | ~93 MB | 250 MB |
| Whisper Tiny (Quantized) | audio | 39M | ~67 MB | 200 MB |
| Whisper Base (Quantized) | audio | 74M | ~121 MB | 300 MB |
| Whisper Small (Quantized) | audio | 244M | ~373 MB | 700 MB |
| Whisper Large-v3 Turbo | audio | 809M | ~1.0 GB | 2.5 GB |
FAQ
Still skeptical? Fair.
Do I need an account or subscription?
No. Everything runs in your browser — no sign-up or billing for core tools. Optional web search uses your own Tavily or Brave API key.
Where do the models come from?
Open-source models distributed via Hugging Face (SmolLM2, Qwen, Whisper, Kokoro, SD-Turbo, SDXL-Lightning and more). Each one is downloaded once, then cached locally on your device.
Is my data ever sent to a server?
AI inference runs on your device, so prompts, files, images, and recordings are not sent to WKO AI. Site assets, model downloads, cookie-free performance analytics, and optional providers use network requests as described in our Privacy Policy.
What hardware do I need?
A WebGPU-capable GPU (Chrome or Edge) gives the fastest results. Most tools automatically fall back to WebAssembly on the CPU; image generation and the largest chat models require WebGPU. Inference remains local.
Why is the first run slow?
The model weights have to be downloaded once. After that, models load from local cache and may work offline when all required assets are cached.
Storage
Manage downloaded models
Model weights are cached in your browser to reduce repeat downloads. Test each tool before going offline. Removing a model frees that space — it downloads again the next time you use it.
Checking local storage…
- Scanning browser storage…
Unlimited local AI. No account required.
AI content is processed on your device. Model downloads and providers you choose may use the network, as explained in our Privacy Policy.