- 4-bit costs about 3 points, 3-bit about 20 or more
Context
Running a language model for Amharic in practice means a small model, heavy quantization and consumer hardware. Data residency, cost and unreliable connectivity push practitioners that way, and it is also the corner of the design space with the least published evidence. Two effects are established separately: small models quantize differently from large ones, and non-Latin scripts suffer more under quantization, although that multilingual claim is contested by a study that found no disproportionate harm on a large model. Nobody had measured the compound for Amharic.
So I asked a practitioner's question. On one consumer GPU, how small a model and how aggressive a quantization can I choose for an Amharic task before accuracy falls off a cliff, what explains the drop, and does cheap fine-tuning buy the loss back? This is independent research on public datasets only, not affiliated with any employer. The paper, "How small, how quantized? Tokenizer, size and quantization for Amharic SLMs on consumer GPUs", is submitted to a 2026 workshop, in review.
Constraints
- The hardware is the subject, not a proxy for it: everything runs on a single 4 GB laptop GPU.
- Public benchmarks only: MasakhaNEWS for topic classification and AfriSenti for sentiment, each paired with an English counterpart so that the script, not the instruction, is the variable.
- Statistics first: fixed seeds, deterministic decoding, bootstrap confidence intervals, identical prompts across languages, and reruns verified byte-identical.
- Connectivity. Full-precision weights beyond a modest size would not download reliably, which limited the controlled arm and the fine-tuning arm to the smallest model. That limit is itself part of the regime under study.
Architecture
The experimental design is the architecture. The matrix covers instruction-tuned checkpoints from four tokenizer families, Qwen 2.5, Gemma 2, Llama 3.2 and Phi-4, with models up to 8B parameters, plus a size ladder inside one family so that scale can be read with the tokenizer held fixed. Precision has two arms. The deployment-realistic arm runs GGUF q8_0, q4_K_M, q3_K_M and q2_K through Ollama, which is what practitioners actually run. The controlled arm quantizes identical weights with bitsandbytes to fp16, int8 and nf4, to check the GGUF result with an independent method.
Two scoped tasks, each paired Amharic and English, give a chance floor that makes degradation readable. The mechanism variable is tokenizer fertility: tokens per word on paired headlines, Amharic against English, per tokenizer. Recovery is QLoRA on the 4-bit smallest model with the Amharic sentiment training split, evaluated on the disjoint test split. A resumable harness writes one row per model, precision, task and language cell, and every table and figure regenerates from those files.
Decisions and tradeoffs
- Scoped classification tasks instead of broad benchmarks. They are the tasks people deploy, they are cheap to run on a small GPU, and their chance floor makes a drop legible. The cost is that the study says nothing about generation or reasoning.
- GGUF through Ollama as the primary arm. It measures what gets deployed, but a quirk of one quantization library could pass for a language effect, which is why the bitsandbytes arm exists.
- English as a paired control with the same instruction. It isolates the script effect. Label sets differ by language on the topic task, which understates the effect rather than inflating it.
- Measuring fertility rather than citing it. The token premium for non-Latin scripts is known; turning it into a per-model predictor of task accuracy is the contribution.
Outcome
The ordering the paper argues is tokenizer first, then model size, then quantization. The first-order driver of Amharic accuracy is the tokenizer: Ge'ez-script fertility varies several-fold across families, and on the topic task it tracks baseline Amharic accuracy almost perfectly across four families, while English accuracy stays flat. A smaller model with a better Ge'ez tokenizer beat larger models with worse ones. Size helps, then saturates early.
Quantization is secondary. In the clearest case, a mid-sized Qwen 2.5 model on Amharic topic classification, 4-bit costs about 3 points, 3-bit about 20 or more, while English degrades gently; a second family shows the same extra Amharic penalty at 3-bit, more mildly. At 2-bit both languages collapse into non-label output, a model break rather than a script effect. The controlled arm confirms, on identical weights, that 4-bit is benign.
QLoRA at the smallest scale did not recover the gap: the tuned model collapsed to predicting a single label. My reading is that the tokenizer had already shredded the Amharic input into fragments a model that small could not use. You cannot fine-tune past a tokenizer that destroys the signal. The guidance follows: choose the tokenizer first, do not quantize Amharic below 4-bit, and prefer a smaller well-tokenized model over a larger poorly-tokenized one.
What I would change
- Run the recovery experiment at larger sizes once the weights can be transferred; whether a mid-sized QLoRA closes the Amharic gap is the open question.
- Add tokenizer families to tighten the correlation; the sentiment result points the same way but is underpowered.
- Extend to generation and reasoning tasks, and to other Ethiopian languages and scripts.
- Add human evaluation, since prior work shows automatic metrics underestimate how much degradation people notice.
Stack and links
Ollama and GGUF for the deployment arm, bitsandbytes and Hugging Face Transformers for the controlled arm, QLoRA for recovery, PyTorch and Python, with MasakhaNEWS and AfriSenti as the public tasks. The repository stays private while the paper is in review; both will be linked once review ends.