juliandClaude Opus 4.8 60cadd0adb Q8 model, confidently-wrong persona, sleek white-on-black UI
- Switch model to Q8_0 (~1.9 GB)
- Add tsjet persona: always replies, plausible-sounding but hilariously wrong
- Bump sampling (temp 0.9 / top-k 60) for playful output; cap replies at 384 tokens
- Redesign UI: monochrome white-on-black, generic IT sans font stack

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-09 20:35:10 +02:00

tsjetpiti

A dead-simple Android chat app that runs an uncensored Qwen3.5-2B model fully on-device via an embedded llama.cpp, wrapped in an extremely barebones WebView UI.

Talk to the tsjet. Hit new tsjet to wipe the conversation and start fresh.


How it works

┌─────────────────────────────────────────┐
│ MainActivity (Kotlin)                    │
│   • full-screen WebView  ── UI ──────────┼──> assets/web/{index.html,app.js,styles.css}
│   • JS bridge "TsjetNative"              │        (the whole chat UI, ~200 lines)
│   • model download + prompt formatting   │
└──────────────┬──────────────────────────┘
               │ JNI (LlamaBridge)
               ▼
┌─────────────────────────────────────────┐
│ cpp/llama-jni.cpp  →  libtsjet.so        │
│   talks to llama.cpp C API (pinned)      │
└─────────────────────────────────────────┘
  • Inference: llama.cpp compiled from source (pinned tag b10333) through the NDK. Pulled automatically at build time via CMake FetchContent — no submodule to init.
  • Model: Qwen3.5-2B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf (~1.2 GB) is downloaded on first launch from Hugging Face into the app's private storage. It is not bundled in the APK.
  • Conversation: the web layer holds the full history and sends it each turn; native rebuilds the ChatML prompt and clears the KV cache before every reply, so "new tsjet" is just: clear JS state + reset cache.

Requirements

  • A build machine with the Android SDK + NDK. Easiest path: open the project in Android Studio (it will offer to install the matching NDK 27.2.12479018 and CMake 3.22.1).
  • A phone: 64-bit ARM (arm64-v8a), Android 8.0+ (minSdk 26), and enough free RAM to hold a ~1.2 GB Q4 model (~34 GB RAM device recommended).
  • Network on first launch to download the model.

Build

./gradlew assembleDebug

The APK lands in app/build/outputs/apk/debug/. Or just Run ▶ from Android Studio.

First build compiles llama.cpp from source, so it takes a while and needs network (CMake fetches the pinned llama.cpp).

Where things live

What Where
Chat UI (HTML/CSS/JS) app/src/main/assets/web/
Android glue + model download app/src/main/java/monster/autisme/tsjetpiti/
Native llama.cpp bridge app/src/main/cpp/llama-jni.cpp
Pinned llama.cpp version app/src/main/cpp/CMakeLists.txt (GIT_TAG b10333)
Model URL / filename ModelDownloader.kt
Generation params (ctx, temp, tokens) MainActivity.kt + llama-jni.cpp

TODO / notes

  • Logo: juli is on it. The header logo is a placeholder 🐟 in assets/web/index.html (#logo), and there's no launcher icon yet — add an android:icon in AndroidManifest.xml + a mipmap when the artwork is ready.
  • The model may emit <think>…</think> blocks; the UI strips them from the display and from history (stripThink in app.js).
  • Bumping the llama.cpp tag? Re-check the C API calls in llama-jni.cpp against that tag's include/llama.h — it uses the raw C API directly.
  • Not yet compiled end-to-end in CI; first real build is the acid test for the native layer.
S
Description
Barebones chat app for a local Ollama model (Qwen3.5-2B) — the tsjet.
Readme
484 KiB
2026-08-09 20:44:13 +00:00
Languages
Kotlin 39.9%
C++ 20.9%
CSS 12.1%
JavaScript 11.4%
Python 7.8%
Other 7.9%