Teaching Tower Thai: our plan for a TowerInstruct LoRA
We are building a small LoRA on Unbabel TowerInstruct-7B so English→Thai machine translation can improve on our own GPU—without burning frontier tokens on every paragraph.
We are building a small LoRA on Unbabel TowerInstruct-7B so English→Thai machine translation can improve on our own GPU—without burning frontier tokens on every paragraph.
We handed a real multi-file React Native feature to a local Qwen2.5-Coder 14B coding agent on our Gaia GPU. It reported success. Git stayed clean. Here is the failure mode — and how we route AI engineering work now.
Hybrid Grok desk: frontier judgment, gaia-tower for EN→ZH/NL. 59 articles translated locally; ~$12 saved this batch.
After open weights lost to Grok 4.5 on our coding desk, a quad-3090 + EPYC box looked like a desinvestment. Hybrid is a first-class Grok Build feature: keep Grok on the API for judgment, offload bulk turns to local inference — with cost calculus.
By 11 August 2031 the main path for Chinese frontier AI does not depend on smuggled NVIDIA boards and Western HBM. The race is the memory wall, packaging, megawatts, and the compiler — not the next gadget roadmap.
Open-weight models are sold as more open than API models. A neural net is billions of opaque floats either way. Open mostly means run it yourself — usually on NVIDIA, sometimes the Ascend bet.
DeepSeek’s new model generation, tuned for Huawei Ascend, is the software half of the NVLink-vs-open-bus fork. How likely is Chinese hardware independence? High for a sovereign inference lane; contested at the deepest manufacturing layers.
Pulling NVLink from RTX 40-class GeForce was not a gamer gift. It reserved multi-GPU fabric for enterprise inference while home labs got PCIe 5.0 ×16. Huawei’s open UnifiedBus 2.0 specs point the other way: link many NPUs as one machine.
Published June 2026 · Infrastructure notes from 3DN This is not a political manifesto and it is not written to shock anyone. It is a note about electricity — the kind of boring constraint that decides whether an AI strategy … Continued
Four RTX 3090s, 256 GB RAM, and GLM 5.2 running scheduled server-admin jobs — inference and operations that stay on premises.