- ModelDownloader now pulls the Q6_K quant - README: reflect Q6_K sizes + note prebuilt APK on Releases Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
tsjetpiti
A dead-simple Android chat app that runs an uncensored Qwen3.5-2B model fully on-device via an embedded llama.cpp, wrapped in an extremely barebones WebView UI.
tsjet has a deliberate personality: it always answers, sounds confident and plausible, and is hilariously, purposely wrong. Hit new tsjet to wipe the conversation and start fresh.
How it works
┌─────────────────────────────────────────┐
│ MainActivity (Kotlin) │
│ • full-screen WebView ── UI ──────────┼──> assets/web/{index.html,app.js,styles.css}
│ • JS bridge "TsjetNative" │ (the whole chat UI, ~200 lines)
│ • model download + prompt formatting │
└──────────────┬──────────────────────────┘
│ JNI (LlamaBridge)
▼
┌─────────────────────────────────────────┐
│ cpp/llama-jni.cpp → libtsjet.so │
│ talks to llama.cpp C API (pinned) │
└─────────────────────────────────────────┘
- Inference: llama.cpp compiled from source (pinned tag
b10333) through the NDK. Pulled automatically at build time via CMakeFetchContent— no submodule to init. - Model:
Qwen3.5-2B-Uncensored-HauhauCS-Aggressive-Q6_K.gguf(~1.5 GB) is downloaded on first launch from Hugging Face into the app's private storage. It is not bundled in the APK. (Swap the quant inModelDownloader.kt.) - Persona: set via
SYSTEM_PROMPTinMainActivity.kt; sampling is a little hot (temp 0.9) for playful answers. The model is a Qwen3 "thinking" model, so the prompt disables reasoning (empty<think></think>prefill +/no_think) to keep replies fast and punchy instead of burning the token budget on hidden thoughts. - Conversation: the web layer holds the full history and sends it each turn; native rebuilds the ChatML prompt and clears the KV cache before every reply, so "new tsjet" is just: clear JS state + reset cache.
Requirements
- A build machine with the Android SDK + NDK. Easiest path: open the project
in Android Studio (it will offer to install the matching NDK
27.2.12479018and CMake3.22.1). - A phone: 64-bit ARM (
arm64-v8a), Android 8.0+ (minSdk 26), and enough free RAM to hold a ~1.5 GB Q6 model (~4 GB RAM device recommended). - Network on first launch to download the model.
Build
./gradlew assembleDebug
The APK lands in app/build/outputs/apk/debug/. Or just Run ▶ from Android Studio.
First build compiles llama.cpp from source, so it takes a while and needs network (CMake fetches the pinned llama.cpp).
Install (sideload)
A prebuilt app-debug.apk is attached to every Release.
The debug APK is signed with the debug key, so you can sideload it directly:
copy app-debug.apk to the phone, enable "install unknown apps" for your file
manager, and tap it. First launch downloads the ~1.5 GB model over Wi-Fi.
Where things live
| What | Where |
|---|---|
| Chat UI (HTML/CSS/JS) | app/src/main/assets/web/ |
| Android glue + model download | app/src/main/java/monster/autisme/tsjetpiti/ |
| Native llama.cpp bridge | app/src/main/cpp/llama-jni.cpp |
| Pinned llama.cpp version | app/src/main/cpp/CMakeLists.txt (GIT_TAG b10333) |
| Model URL / filename | ModelDownloader.kt |
| Generation params (ctx, temp, tokens) | MainActivity.kt + llama-jni.cpp |
TODO / notes
- Branding: launcher icon (adaptive + legacy densities) is generated from
piti-icon.pngviapython3 tools/gen_launcher_icons.py; the header wordmark ispiti-logo.png(copied toassets/web/logo.png). Source art lives at the repo root. - The model may emit
<think>…</think>blocks; the UI strips them from the display and from history (stripThinkinapp.js). - Bumping the llama.cpp tag? Re-check the C API calls in
llama-jni.cppagainst that tag'sinclude/llama.h— it uses the raw C API directly. - Native is always built optimized. AGP compiles the debug variant's C/C++ at
-O0by default, which makes llama.cpp ~10x too slow;CMakeLists.txtforcesRelease/-O3regardless of variant. Native libs are also 16 KB page-aligned. - Verified on-device (Pixel 7, Android, 4 KB pages): downloads the model, loads
it (
n_ctx=4096, 6 threads), streams a reply, and "new tsjet" clears the chat../gradlew assembleDebug→ ~13 MBarm64-v8aAPK (model downloads on first launch).