Qwen 3.8 vs Qwen 3.5: What Our Blog Experiment Taught Us

Qwen 3.8 has been one of the most impressive open-source models I’ve personally tested. The last model release that left me similarly excited was DeepSeek. That is my experience, not a benchmark ranking—but it gave us a practical question to investigate.
Could we keep the larger model available for coding and let a smaller local model handle blog writing or translation? We tried it with our own technical documentation, saved the outputs, and checked what the models actually said.
The result was more interesting than simply declaring a winner.
What we actually compared
Our local models were Qwen 3.8 27B and Qwen 3.5 9B, both using Q4_K_M quantization through Ollama. Their metadata reported approximately 27.3 billion and 9.7 billion parameters. They ran on different servers.
That matters: model generation, model size, and hardware all changed. This experiment compares two practical setups. It cannot isolate the effect of the version number.
The writing task covered a real technical story: preparing an older desktop for local inference. Its facts contained distinctions that were easy to lose. A GPU fault remained unresolved. Some configuration was owner-reported rather than independently inspected. A smoke test was not a reliability qualification. A generated code example was inspected, not executed.
Those qualifications were part of the story, not optional wording.
Fluent writing got ahead of the evidence
Some early Qwen 3.5 drafts added stability or production-readiness conclusions the documents did not support. The sentences sounded plausible, but the evidence did not establish those outcomes.

Fluent writing got ahead of the evidence
Then Qwen 3.8 made a similar mistake. With the old prose instructions, one draft claimed that a single smoke test verified stability.
Switching models had not solved the whole problem.
We expanded a thin six-fact brief into 17 source-checked facts and changed the instructions to factual reporting. During that exploratory phase, several things changed together, so we could not attribute improvement to one setting.
We then ran a matched comparison: the same facts, prompt, JSON schema, requested 8,192-token context, sampling settings, and two seeds per model. Writing temperature was 0.3, with thinking disabled. An audit confirmed that each paired request differed only in the model name.
The four resulting bodies were 473 and 483 words for Qwen 3.8, and 492 and 503 words for Qwen 3.5. All met our structural limits. An AI-assisted reviewer checked copies with model labels hidden and found all four main bodies materially grounded.
There was no strong factual winner in that small matched rerun.
Both still needed editing. Qwen 3.8 added punctuation inside an exact output quotation and used a title that could confuse single-GPU inference with a physically single-GPU machine. Qwen 3.5 used ambiguous “memory available” wording and copied instructions into image alt text.
The smaller model’s improvement is important evidence. Calling it incapable of writing would ignore what happened after we improved the setup.
Translation exposed a different set of mistakes
We translated one 507-word English article, plus its title, excerpt and SEO fields, into Spanish, French, Brazilian Portuguese, Italian, and Mandarin written in Simplified Chinese.
Each model ran once per language with basic translation instructions, then again with added terminology guidance: ten translations per model, twenty outputs total. Requests used temperature zero, an 8,192-token context, and thinking disabled. Again, paired requests matched except for model name.
All twenty completed as valid JSON. Several still changed the meaning.
| First-pass example | Qwen 3.5 | Qwen 3.8 | | --- | --- | --- | | Spanish: “cloud fallback” | Wording suggesting a cloud outage | Wording that could mean cloud backup | | French: binary memory units | Changed GiB/MiB to Go/Mo | Preserved binary units as Gio/Mio | | Chinese: “out-of-memory” | Rendered as memory overflow | Rendered as insufficient memory |
Qwen 3.8 also preserved an English test-output quotation that Qwen 3.5 translated despite instructions to keep it unchanged. In French, the larger model correctly said the system “returned” the requested text; the smaller model used a verb meaning “retained.”
Those were useful differences. They were also specific findings from one technical article, not a general language ranking.
Better instructions helped, but introduced surprises
We added guidance for technical units, cloud fallback, insufficient memory, literal quotations, and the meaning of a “warm” sample.
That corrected the main Qwen 3.5 unit, cloud-fallback, and Chinese memory-error and quotation problems. Sometimes it avoided an error by retaining English jargon, which still needed copyediting. Its French “retained” versus “returned” mistake remained.
Qwen 3.8 was not immune to regressions. In the terminology-guided French pass, container runtime registration became wording meaning “execution time recorded.” Portuguese runtime wording also became ambiguous. Its Spanish cloud-fallback translation improved.
A more explicit prompt helped several sentences while making another worse. Checking only the errors from the previous run would have missed that.
Thinking was disabled throughout these matched tests. That was enough to complete the tasks, but we did not compare thinking-on translations. We therefore cannot claim that reasoning mode is never useful—or that disabling it guarantees faithful translation.
What we’re choosing—and what remains unproven
I still prefer Qwen 3.8 for our English articles, based on the broader experience that led to this experiment. Qwen 3.5 remains worth testing for constrained translation drafts, with terminology guidance and review.
The evidence supports that practical choice without requiring us to dismiss the smaller model. Configuration mattered substantially, and the matched writing results narrowed the apparent gap.
These were small, AI-assisted checks: two matched writing drafts per model and translations of one technical article. We did not use a professional native-speaker panel or establish general accuracy rates. Different hardware also prevents treating local response times as an inherent model-speed advantage.
My enthusiasm for Qwen 3.8 survives the experiment. So does the requirement to check its work. For this workflow, preserving a qualification, a unit, or the meaning of a technical term matters as much as producing a polished sentence.