Skip to content
WKO AI

23 open models · WebGPU accelerated · No ad tracking

Free AI Chat, Images & Voice — 100% In Your Browser

No signup. Your files never leave your device. Runs locally after setup. Prepare and test before going offline.

  • WebGPU accelerated
  • WASM fallback
  • 23 open models
  • Local processing

Choose a tool

Start where you are

What would you like to do?

What's inside

Three studios, one browser tab

Chat

  • Qwen 3.5 & more

    Switch models mid-conversation

  • Thinking mode

    Watch compatible models reason before answering

  • Chat with your files

    PDF & text knowledge via local RAG

  • Voice in, voice out

    Local transcription input, Kokoro read-aloud replies

Open Chat

Image

  • Text-to-image

    SD-Turbo fast drafts or SDXL-Lightning quality

  • Image to text

    Extract, copy, and ask about printed text

  • Depth maps

    Depth Anything 3 scene structure

  • Upscale

    Swin2SR 2×/4× super-resolution

  • Remove background

    BiRefNet one-click cutouts

Open Image

Audio

  • Neural TTS

    Kokoro-82M with 27 voices & 3 speeds

  • Transcribe

    Moonshine English & multilingual Whisper

  • Clip history

    Playback, timestamps, and subtitle exports

  • Record & import

    Live mic capture with level meter

Open Audio

How it works

Local in three steps

  1. 01

    Pick a tool

    Chat, Image, or Audio. No account, no setup, nothing to configure.

  2. 02

    Model downloads once

    Open weights come straight from Hugging Face and are cached in your browser.

  3. 03

    Runs locally & privately

    Inference happens on your GPU or CPU. AI content stays on your device unless you activate an external provider.

Private by design

Runs on your device. Proof, live.

Session readout
0 / 23

Models cached

Cached model files — processing runs on your device

Checking device…

Model weights download once from Hugging Face, then inference runs locally. Test cached tools before disconnecting.

Browser support
  • Chrome / Edge WebGPU
  • Firefox WASM fallback
  • Safari WASM fallback

WebGPU unlocks GPU-accelerated inference. Firefox and Safari can run most tools through the WASM CPU path — image generation (SD-Turbo, SDXL-Lightning) and the largest chat models require WebGPU.

AI inference is local. · Site assets, cookie-free performance analytics, model downloads, and providers you enable may use the network.
Full model matrix 23 models
Model Area Parameters Download VRAM
SmolLM2 360M chat 360M (q4f16_1) ~220 MB 400 MB
LFM2.5 350M chat 350M (q4f16) ~210 MB 450 MB
Qwen 3.5 0.8B chat 0.8B (q4f16) 430 MB ~1.6 GB
Qwen 3.5 2B chat 2.3B (q4f16) 1.1 GB ~2.2 GB
Qwen 3.5 4B chat 4B (q4f16) 2.3 GB 3.9 GB
Qwen 3.5 9B chat 9B (q4f16) 4.8 GB 6.4 GB
SD-Turbo image Single-step Turbo 2.3 GB 1.2 GB
SDXL-Lightning image 2.6B UNet (INT4) ~3.6 GB 6.0 GB
Depth Anything 3 Small image ~26M 105 MB 300 MB
Depth Anything 3 Base image ~103M ~413 MB 800 MB
Swin2SR Lightweight 2x image 1.01M · fp32 8.08 MB 250 MB
Real-ESRGAN General x4v3 image 1.21M · fp32 4.87 MB 300 MB
Swin2SR Classical 4x image 12.2M · q8 21.6 MB (q8) 500 MB
Fast image BiRefNet Lite (512) ~99 MB 300 MB
Quality image BEN2 Base ~219 MB 600 MB
PP-OCRv5 Mobile image Det 4.7M + Rec 15.8M ~21.3 MB 150 MB
Kokoro 82M audio 82M 86 MB 200 MB
Moonshine Tiny audio 27M ~53 MB 150 MB
Moonshine Base audio 61M ~93 MB 250 MB
Whisper Tiny (Quantized) audio 39M ~67 MB 200 MB
Whisper Base (Quantized) audio 74M ~121 MB 300 MB
Whisper Small (Quantized) audio 244M ~373 MB 700 MB
Whisper Large-v3 Turbo audio 809M ~1.0 GB 2.5 GB

FAQ

Still skeptical? Fair.

Do I need an account or subscription?

No. Everything runs in your browser — no sign-up or billing for core tools. Optional web search uses your own Tavily or Brave API key.

Where do the models come from?

Open-source models distributed via Hugging Face (SmolLM2, Qwen, Whisper, Kokoro, SD-Turbo, SDXL-Lightning and more). Each one is downloaded once, then cached locally on your device.

Is my data ever sent to a server?

AI inference runs on your device, so prompts, files, images, and recordings are not sent to WKO AI. Site assets, model downloads, cookie-free performance analytics, and optional providers use network requests as described in our Privacy Policy.

What hardware do I need?

A WebGPU-capable GPU (Chrome or Edge) gives the fastest results. Most tools automatically fall back to WebAssembly on the CPU; image generation and the largest chat models require WebGPU. Inference remains local.

Why is the first run slow?

The model weights have to be downloaded once. After that, models load from local cache and may work offline when all required assets are cached.

Storage

Manage downloaded models

Model weights are cached in your browser to reduce repeat downloads. Test each tool before going offline. Removing a model frees that space — it downloads again the next time you use it.

Checking local storage…

  • Scanning browser storage…

Unlimited local AI. No account required.

AI content is processed on your device. Model downloads and providers you choose may use the network, as explained in our Privacy Policy.