Apache-2.0updated 19d ago
Whatever is loaded here is a 0.5B–4B class model — the shipped default is Gemma 4 E4B, but LITERTLMPLUGINMODEL and --model both change that, so check with --check rather than assuming. It is not a small Claude. Prompting habits that work on a frontier model — layered instructions, implied reasoning, "think about X then do Y" — degrade its output sharply.
What can you do with Litertlm Prompting?
name: litertlm-prompting description: How to prompt a small on-device model so it performs. Load this before composing a prompt for the local model — habits that work on frontier models actively hurt here.
Prompting a small local model
Whatever is loaded here is a 0.5B–4B class model — the shipped default is Gemma 4 E4B, but
LITERT_LM_PLUGIN_MODEL and --model both change that, so check with --check rather than
assuming. It is not a small Claude. Prompting habits that work on a frontier model — layered
instructions, implied reasoning, "think about X then do Y" — degrade its output sharply.
One family-specific trap worth knowing: some models in this range are reasoning variants
(names carrying Thinking, R1, VibeThinker, and Qwen3 base models) and emit a
<think>…</think> block before answering. That burns your token budget before a single useful
word. Prefer an Instruct variant when one exists, and raise --max-tokens if you must use a
thinking model, or the answer gets truncated mid-reasoning.
Most "this model is useless" conclusions are prompt problems. Fix the prompt before concluding the model can't do it.
(For how to report what it returns, see the litertlm-usage skill. This skill is about
getting a good answer; that one is about not overselling it.)
The rules that matter most
1. One task per call
Compound instructions are where small models fall apart first. Each additional clause competes for attention and the later ones get dropped or half-served.
BAD Review this function, suggest fixes, and rewrite it with tests.
GOOD List the defects in this function. One per line.
If you need three things, make three calls. They cost nothing.
2. State the output shape
Left unspecified, it will produce prose padding. Given a shape, it fills the shape.
BAD What's wrong with this regex?
GOOD Explain this regex in exactly three bullet points: what it matches,
what it rejects, one edge case.
3. Everything it needs must be in the prompt
It has no file access and no repository awareness. Any reference to something not pasted or piped in will be answered by invention, fluently. Paste the code. Pipe the diff.
4. Do not ask for reasoning chains
"Think step by step before answering" reliably makes it worse — it generates plausible-looking reasoning that doesn't constrain the final answer, and burns context doing it. Ask for the answer. If you want the reasoning, ask for it as a separate, second call.
5. Prefer closed questions to open ones
It is markedly better at judging than at generating.
WEAKER How should I structure this module?
STRONGER Here are two structures, A and B. Which has fewer failure modes, and why?
6. Use --system for the role, the prompt for the task
Keep the role stable and the task specific. Roles that constrain output length work well:
--system "You are a terse technical reviewer. Answer in under 100 words. \
If you are unsure, say so rather than guessing."
The "say so rather than guessing" clause measurably reduces confident invention. Worth including whenever the answer might be outside what you supplied.
7. Keep input short
A long prompt fails, and how it fails depends on the runtime. Measured on litert-lm 0.14.0
serving qwen3-4b-instruct, it breaks the HTTP response rather than answering — the caller
sees a transport error, not a short reply. Other runtimes truncate silently instead and answer
from what fit, which is worse because nothing marks the reply as partial. Assume either.
The ceiling is lower than the model card suggests: serve runs at a fixed
max_num_tokens=4096 regardless of the model's native context, which came to roughly
6.5–8 KiB of diff text depending on how it tokenises. Narrow a large diff to one file. Excerpt
the relevant function rather than pasting the module. If the input feels big, it is.
8. Set --max-tokens to what you actually want
It fills the budget it is given. A 900-token budget on a yes/no question produces 900 tokens of justification.
Where it is genuinely strong
- Explaining a self-contained snippet, regex, error string, or config
- Rephrasing, summarising, tightening prose you supply
- Closed comparisons between options you spell out
- Naming things, drafting boilerplate
- Extracting structure from unstructured text you paste in
Where it will fail regardless of prompting
- Anything requiring repository knowledge
- Multi-step reasoning where step three depends on step one
- Precise arithmetic
- Recalling specific API signatures or version details — it will confabulate them fluently
For those, use Claude directly. No amount of prompt engineering closes a capability gap.
Multi-turn and tool calling
Both work, with caveats:
- Multi-turn requires you to resend the whole history in
messages; the client does not do this for you. Each call is otherwise independent. - Tool calling returns correctly-shaped
tool_calls, but nothing executes them — the plugin surfaces the request and stops. Keep tool schemas small; a large tool array crowds out the actual prompt.
Install
Add Litertlm Prompting to your client. Pick the one you use.
npx skills add kurtvalcorza/litert-lm-plugin-ccInstalls every skill in the repository, then prompts for which to keep.
/plugin marketplace add kurtvalcorza/litert-lm-plugin-ccAdds the repository as a plugin marketplace; install individual plugins with `/plugin install`.
git clone https://github.com/kurtvalcorza/litert-lm-plugin-cc
cp -r plugins/litertlm/skills/litertlm-prompting ~/.claude/skills/A skill is a plain directory. Copy it into `.claude/skills/` in a project or in your home directory.
Score
79 / 100
Good