Pixel-art fox mascot standing idle

Android · Local-first AI

Private AI that runs on your phone.

Run GGUF and LiteRT-LM models 100% offline on your Android phone. No internet, no account, no data leaves your device. Optional cloud API fallback when you choose.

No account · No cloud by default

Get Snor

Install on Android

Download the APK and sideload it, or install from your file manager. Allow “install from unknown sources” when prompted.

Download APK v1.0.5

Signed release build · ~82.5 MB · ARM64 (arm64-v8a) · armeabi-v7a · x86_64

What it does

Everything on-device. Anything off-device, on your terms.

Pixel-art fox dozing with Zs above its head

Local inference

Download and run GGUF models with GPU acceleration via Vulkan, or LiteRT-LM .litertlm models on CPU, GPU, or NPU. No internet required after download.

Any LiteRT-LM model can also serve your local network over an OpenAI-compatible /v1/chat/completions endpoint, with an optional ngrok or Cloudflare tunnel when you need reach from outside.

Pixel-art fox sitting calmly while a cloud request runs

Cloud fallback

Seamlessly switch to OpenAI, Anthropic, Google Gemini, DeepSeek, OpenRouter, or NVIDIA NIM when you need more power. You choose where inference runs.

OpenAI Anthropic Google Gemini DeepSeek OpenRouter NVIDIA NIM
Pixel-art fox holding an open card in a multimodal chat

Multimodal chat

Send text and images in conversations. Vision works with local models (Qwen2-VL) and cloud providers alike.

Pixel-art fox mascot standing idle

Persistent sessions

Chats, tasks, and settings live in local storage via Hive. Nothing leaves your device unless you explicitly pick cloud mode.

Pixel-art fox with raised paws during first-launch auto-configuration

Smart auto-config

On first launch, Snor reads your device’s RAM and sets optimal context size and token limits automatically.

Pixel-art fox sitting and waving

Task workflows

A dedicated task view for structured AI-assisted workflows, alongside free-form chat.

How it works

One message, two engines. You pick the route.

1

Type a message

Prompts start from a single input box and stream back into the same conversation.

Local

GGUF or LiteRT-LM, on-device

llama.cpp with Vulkan, or LiteRT-LM, after a one-time download. No internet required.

Cloud

Cloud model API

OpenAI, Anthropic, Gemini, DeepSeek, OpenRouter, or NVIDIA NIM. Used only when you pick a provider.

2

Response streams back

Answers land in the same chat — local data stays on-device, cloud calls go only to the endpoint you chose.

Privacy

Nothing leaves your device.

Local inference and the optional LAN server run on-device. Chats are stored locally, models are local, and nothing is sent anywhere without an explicit choice.

Even in cloud mode, API keys are kept in local storage and transmitted only to the provider’s endpoint — never to Snor, never to a third party.

FAQ

Common questions

Does Snor work without internet?

Yes, Snor runs GGUF and LiteRT-LM models completely on-device. No internet connection is required for local inference after the initial app and model download.

What GGUF models are supported?

Snor supports GGUF models with Vulkan GPU acceleration. Tested models include Phi-3, TinyLlama, Gemma 2B, and various Llama 3 variants in Q4_K_M and Q5_K_M quantization levels.

Is my data private with Snor?

Absolutely. Your chats, settings, and models remain on your device unless you explicitly choose to use cloud API fallback. Even in cloud mode, API keys are stored locally and transmitted only to the chosen provider.

How do I install Snor?

Download the APK from snor.site and open it on any Android 8.0 or newer device, then allow installs from your file manager when prompted. Pick the build that matches your device ABI: ARM64, armeabi-v7a or x86_64.

Does Snor send my prompts to a server?

Not unless you turn on cloud mode yourself. Local inference runs entirely on your device. In cloud mode, requests go directly from your device to the provider you selected, and your API key is stored locally.

Can I use my phone as a local AI server?

Yes. Any LiteRT-LM model loaded in Snor can expose an OpenAI-compatible /v1/chat/completions endpoint on your local network, so other apps and tools on your Wi-Fi can use it. An ngrok or Cloudflare tunnel is optional if you need access from outside your network.