Teaching Tower Thai: our plan for a TowerInstruct LoRA
We are building a small LoRA on Unbabel TowerInstruct-7B so English→Thai machine translation can improve on our own GPU—without burning frontier tokens on every paragraph.
We are building a small LoRA on Unbabel TowerInstruct-7B so English→Thai machine translation can improve on our own GPU—without burning frontier tokens on every paragraph.
We handed a real multi-file React Native feature to a local Qwen2.5-Coder 14B coding agent on our Gaia GPU. It reported success. Git stayed clean. Here is the failure mode — and how we route AI engineering work now.
Hybrid Grok desk: frontier judgment, gaia-tower for EN→ZH/NL. 59 articles translated locally; ~$12 saved this batch.
After open weights lost to Grok 4.5 on our coding desk, a quad-3090 + EPYC box looked like a desinvestment. Hybrid is a first-class Grok Build feature: keep Grok on the API for judgment, offload bulk turns to local inference — with cost calculus.
Open-weight models are sold as more open than API models. A neural net is billions of opaque floats either way. Open mostly means run it yourself — usually on NVIDIA, sometimes the Ascend bet.
DeepSeek’s new model generation, tuned for Huawei Ascend, is the software half of the NVLink-vs-open-bus fork. How likely is Chinese hardware independence? High for a sovereign inference lane; contested at the deepest manufacturing layers.
If you run coding agents day in and day out, you eventually notice something odd. Some weeks the bill looks almost boring. Then a few sharp spikes appear — usually on days when someone reopened an old chat on another … Continued