On-device fixes: optimized native build, 16 KB align, disable thinking

Validated on a Pixel 7 (adb). Key fixes:
- Force Release/-O3 for native even in the debug variant (AGP defaulted to -O0,
  making llama.cpp ~10x too slow / effectively unusable).
- 16 KB-align all native LOAD segments (Play requirement; 16 KB-page devices).
- Disable Qwen3 thinking (empty <think></think> prefill + /no_think) so replies
  are fast and punchy instead of spending the whole budget on hidden reasoning.
- Cap replies at 220 tokens.

Verified: model downloads, loads (n_ctx=4096, 6 threads), streams an on-persona
reply with emoji intact, and "new tsjet" clears the conversation.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
2026-08-09 22:06:30 +02:00
co-authored by Claude Opus 4.8
parent 20a16c9faf
commit b06a0349a9
4 changed files with 30 additions and 6 deletions
+3
View File
@@ -24,7 +24,10 @@ android {
}
externalNativeBuild {
cmake {
// Force an optimized native build even for the debug APK — otherwise
// AGP compiles ggml/llama at -O0 and inference is ~10x too slow.
arguments += "-DANDROID_STL=c++_shared"
arguments += "-DCMAKE_BUILD_TYPE=Release"
cppFlags += "-std=c++17"
}
}