FluentAI v1.3.2 — Privacy-First AI & RAG

Your AI. Your Device. Your Silicon.

Chat with Llama, Gemma 4, DeepSeek, Mistral, and 100+ models locally. Accelerated by Snapdragon NPU, Apple Silicon MLX, or GGUF runtimes. Completely offline. Completely free. Zero data collection.

FluentAI Chat Interface on Android
What’s new in v1.3

Latest Product Updates

Models

Gemma 4 (E2B + E4B)

Apache 2.0, 128K context. New SoTA local class. Support for Gemma 4n with MTP speculative decoding — up to 2× faster on mobile GPU.

Hardware

NPU on Snapdragon

QNN delegate via Play Feature Delivery. SoC-aware backend selection (QNN → GPU → CPU) for lower battery drain and NPU speeds.

Platform

MLX on Apple Silicon

Real inference on macOS & iOS 18+ A17 Pro+. Metal-native execution with 1-bit affine quantization, bringing 7B models down to ~1.75 GB.

Agents

On-Device AI Agents

Plan-and-execute agents with task memory, scheduled cron runs, composable skills, and native mobile tools (clipboard, contacts, files).

Why FluentAI?

A fully private AI agent platform that runs models natively on your device without monthly subscriptions or data leaks.

Privacy First

Conversations stay entirely on-device. Zero telemetry, zero tracking, and zero data collection. No cloud server required.

100+ AI Models

Run Llama 3, Gemma 4, DeepSeek, Mistral, Phi, or Qwen locally. Or optionally connect to Claude, GPT-4, and Gemini in the cloud using your own API keys.

Multi-Runtime Engine

Three inference backends (Fllama/GGUF, LiteRT/Android NPU, MLX/Apple Silicon) built-in. The app auto-selects the fastest for your device.

Knowledge Bases

Upload PDFs and text files to query your data locally. Fully private Retrieval-Augmented Generation (RAG) with local semantic search.

Tool Calling & MCP

Built-in tools for search, calculations, and memory. Full Model Context Protocol (MCP) support connects your AI to GitHub, Slack, Notion, and more.

Voice Chat

Speak with your AI naturally using 5 distinct conversation modes: Normal, Interview, Learning, Storytelling, and Translation.

Local OpenAI API Server

FluentAI serves /v1/chat/completions directly on your local network. Other apps can use FluentAI as their offline model backend.

Hugging Face Browser

Search 10,000+ GGUF models directly in-app. Features memory-fitness badges to prevent your phone or tablet from running out of RAM (OOM).

Inference Runtimes

One App. Three Runtimes.

FluentAI embeds multiple inference frameworks to deliver the fastest performance regardless of your hardware configuration.

FLM

FllamaRuntime

GGUF · llama.cpp · All Platforms

  • Gemma 4 architecture backport (MoE 128 experts, ISWA dual-cache)
  • KleidiAI v1.23.0 optimizations (SME2 + Q4_K paths)
  • KV cache TQ4/TQ3 quantization for memory efficiency
  • 16 KB page alignment support for Android 15+
LRT

LiteRTRuntime

Android · NPU / GPU · LiteRT-LM 0.10

  • Snapdragon NPU acceleration via QNN delegate
  • SoC-aware backend selection: QNN → GPU → CPU
  • Play Feature Delivery for modular, bloat-free installs
  • MTP speculative decoding for 1.5–2× faster generation
MLX

MlxRuntime

macOS · iOS 18+ A17 Pro+ · Apple Silicon

  • Native Apple MLX inference on M-series chips and iOS
  • 1-bit affine quantization runs 7B models in ~1.75 GB RAM
  • Metal-native execution with zero Rosetta overhead
  • Multi-file parallel downloads direct from Hugging Face

Supported Platforms & Downloads

Android App

Google Play Store

Full hardware NPU acceleration, GPU/CPU execution, and offline local model downloads.

Download on Android
Desktop Clients

macOS, Windows & Linux

Run MLX models on Apple Silicon macOS, or load local GGUF models via Fllama on Windows and Linux.

Visit FluentAI App

Pricing Plans

Choose the plan that suits you best. Support independent open-source AI development.

Free Tier

Perfect for offline, private local AI execution.

$0 / forever
  • 100+ local AI models support
  • Unlimited offline chat conversations
  • Full voice chat with 5 modes
  • Knowledge bases (Private RAG)
  • Tool Calling & MCP support
Included by default
RECOMMENDED

Premium Upgrade

Unlock customization options and support the product.

$3.49 one-time payment
  • Everything in Free Tier
  • 100% Ad-free experience
  • 9 premium visual themes
  • Private cloud sync dashboard
  • Advanced model temperature/settings
  • PDF conversation exports & analytics
Upgrade in App

Frequently Asked Questions

Is FluentAI really free?

Yes! Running local models on your device is completely free and unlimited. If you decide to connect to cloud providers like Claude, GPT-4, or Gemini, you will just need to supply your own API keys.

Does it work completely offline?

Absolutely. Once you download your chosen GGUF or LiteRT models, all conversation, processing, and inference are executed entirely offline on your device, requiring no internet connection whatsoever.

Which AI models can I run?

You can run over 100 models locally including Llama 3, Gemma 4 (E2B/E4B), DeepSeek, Mistral, Phi, and Qwen via GGUF, LiteRT, or MLX formats. You can also import any GGUF model directly by pasting its Hugging Face URL.

How does the NPU acceleration work?

On compatible Android devices with Snapdragon SoCs, FluentAI uses the LiteRT QNN delegate to route tensor math directly to the Snapdragon NPU. This yields up to 4× speedups over standard CPU execution while using significantly less battery.

What platforms are supported?

Android is available now on Google Play. macOS and iOS 18+ (A17 Pro+) versions support real Apple Silicon MLX inference. Desktop versions for Windows and Linux are also available for download.

Start Chatting Privately Today

No credit card, no sign-up, no monthly fees. Just pure on-device AI running with hardware acceleration. Take back your privacy.