Target Hardware: POCO X7 Pro 5G • 8 GB LPDDR5X • Dimensity 8400-Ultra

On-Device AI Engine Specifications

A complete technical reference for combining lightweight, task-specific models into a unified, offline-first personal assistant on Android.

16 Specialist Models
~4.2 GB Safe RAM Headroom
3 Runtimes Sherpa + llama.cpp + ONNX
100% Offline & Private

Model Technical Cards

Filter by domain or search by model name to inspect parameter counts, memory consumption, and runtime configurations.

Master Technical Matrix

Side-by-side comparison of parameter sizes, on-disk footprints, active memory consumption, and runtime engines.

Model Role Parameters Quantization Disk Size Peak RAM Runtime Engine Hardware Target

Interactive 8 GB RAM Allocation Simulator

Select which models you plan to run simultaneously to test whether your setup remains within safe operating limits or triggers Android's Low Memory Killer.

HyperOS System + UI Baseline Reserved: 3,400 MB

Total RAM Consumption

Total Used: 3,400 MB / 8,000 MB
SAFE (Headroom Available)
System memory is within normal operating parameters.

End-to-End Orchestrated Workflows

Execution flows showing data movement across individual specialized models.

Single App Architecture Blueprint

How the native Android orchestration application is constructed using Kotlin and the Android NDK.

Runtime Layer 1

Sherpa-ONNX (JNI)

Direct C++ bindings handling audio hardware I/O:

  • Continuous Microphone Capture: 16 kHz mono buffer streaming.
  • Silero VAD: Real-time zero-copy chunk evaluation.
  • Whisper Tiny & IndicConformer: Fast CPU-based speech decoding.
  • Piper & Kokoro: Text-to-speech output streamed directly to Android AudioTrack.
Runtime Layer 2

llama.cpp (NDK + Vulkan)

Compiled natively as libllama.so to run GGUF quantized models:

  • Mali-G720 GPU Acceleration: Vulkan backend offloading.
  • Qwen3-1.7B Orchestration: Generates strict JSON schema tool calls.
  • Gemma 3 270M: Ultra-fast text classification and intent triage.
  • Moondream 2 VLM: Multimodal visual reasoning and Q&A.
Runtime Layer 3

ONNX Runtime Mobile

Cross-platform inference engine for vision and memory models:

  • BGE-small-en: 384-dimensional vector embeddings for local RAG.
  • PP-OCR Mobile: Fast text detection and line extraction.
  • YOLO11n: Object detection at 60+ FPS via OpenCL/NNAPI.
  • MobileCLIP-S0: Zero-shot image and text similarity matching.