Local GPU bulk MT: Tower-Plus on a 3090 cuts frontier token cost for 3DN and PolitiCap

posted in: Uncategorized | 0

This afternoon we finished a full language wave without burning a frontier API bill on the bulk work: 22 English 3DN posts into Chinese, and 37 PolitiCap desk notes since 21 August into Dutch. The machine that did the cranking sits on the LAN — an NVIDIA RTX 3090 under the desk name Gaia — running Unbabel Tower-Plus-9B through Ollama as the Grok Build pin gaia-tower.

The hybrid desk in one sentence

Grok 4.5 (frontier) still owns judgment, tone, SEO weave, and “is this publishable?”. Local GPU owns mechanical copies: explore/plan workers on Qwen 2.5 Coder 14B (gaia-gpu), and article machine translation on Tower-Plus (gaia-tower). Parents stay expensive; bulk tokens stay on VRAM we already own.

Cost comparison frontier vs local Tower
One batch: estimated frontier API cost vs electricity on the 3090.

What actually ran

  • 3dn.nl — Polylang gained zh; every published English post now has a linked Chinese sibling.
  • politicap.eu — Polylang gained nl; every English wire from 21 Aug onward has a Dutch sibling. Switcher order is EN | NL | 中文 | TH.
  • HTML structure, ticker links, and brand symbols stayed intact; only human-readable text nodes moved.
Batch counts ZH and NL
Tower-Plus throughput for this run: 22 ZH + 37 NL articles.

Earnings math (this batch)

Rough but honest order-of-magnitude for ~400k characters of article HTML (titles + chunked bodies, prompt overhead included):

Lane Estimate
Frontier API if the same bulk MT ran on Grok-class pricing (~150k input + ~150k output tokens with chunk overhead) ≈ $8–15 (we book $12 mid)
Local Tower-Plus on RTX 3090 (~25–40 min wall, ~300–350 W GPU draw) ≈ 0.15 kWh · ~$0.02–0.04 electricity
Saved this afternoon ≈ $12 on one wave

That is not “free AI”. It is token cost moved onto compute we already pay for as infrastructure. Repeat every publish wave and the annual number is the interesting one: a desk that ships bilingual/trilingual posts weekly can keep four-figure frontier spend off the credit card while the 3090 stays warm for OCR (Typhoon) and coding workers (Qwen) the rest of the day.

Grafana: hybrid Gaia GPU (live panels)

Screenshots from the live fleet dashboards via Grafana’s /render API (remote image renderer on the build host). Source boards: Grok Build · hybrid Gaia GPU and Grok CLI · queries & cost.

Grafana token rate by model
Token rate by model (parent Grok API vs Gaia local) — last 24h.
Grafana tokens in range by model
Tokens in range by model — last 7 days.
Grafana local share of tokens
Local share of Grok Build OTEL tokens (7d). Bulk Tower MT hits Ollama directly, so desk OTEL share stays low while Gaia GPU gauges still show the load.
Grafana Gaia GPU util and power
Gaia RTX 3090 — GPU utilization and power (24h).
Grafana Gaia VRAM
Gaia VRAM (Ollama loaded + nvidia used) — 24h.
Grafana estimated cost rate
Estimated cost rate by token type (Grok CLI cost board, 7d).

Observability

Gaia’s Ollama exporter is on the fleet Prometheus path (gaia_ollama_up, model inventory, VRAM gauges). The hybrid Grok Build path emits OTEL into the same stack; Grafana’s hybrid desk view is where we watch local vs frontier share once traffic is flowing. After a power blip we also locked Ollama into a boot task so LAN inference does not depend on a desktop session staying open.

Why this fits 3DN

Digital sovereignty is not only “which cloud”. It is whether bulk inference for our own properties has to leave the building. Managed hosting and coding-agent desks already live on our iron; bulk MT now sits next to them. Frontier Grok remains the sharp pen. The GPU is the printing press.

Fintech spine for the family: cheaper publish ops means more desk capacity for products that move money — DutchBud bank, closed-loop credits, and the civic market around PolitiCap — without pretending we are a licensed open-banking provider.

Leave a Reply

Your email address will not be published. Required fields are marked *