เดลต้า ไม่ใช่แม่น้ำ: รูเตอร์เป็นหน่วย, การนอนหลับเป็น HA, อินเทอร์เน็ตเป็นแผนที่

posted in: Uncategorized | 0

โมเดล Frontier ล้มเหลวในการทดสอบการดำเนินงานสามประการพร้อมกัน: พวกมันไม่สามารถเรียนรู้ทักษะใหม่ได้โดยไม่เสี่ยงต่อทักษะเดิม พวกมันไม่มีกะกลางคืน และ serving weights เป็นสำเนาเพียงชุดเดียว การเรียนรู้แบบเพิ่มขึ้น (Incremental learning), ความยืดหยุ่น (plasticity) และ การดูแลรักษา (maintenance) เป็นปัญหาเดียวกันที่มีสามชื่อ คำตอบปกติคือแม่น้ำที่ใหญ่ขึ้น — HBM ที่มากขึ้น, tokens ที่มากขึ้น, mesh เดิม

3DN ต้องการ delta.

ภาษาคือรถเมล์ — tokens คือสกุลเงิน

งานของ LLM ไม่ใช่แค่ “พูด” หลังจากการ pretraining มันมี token space ที่ใช้ร่วมกัน, next-token control และ pattern ของโลกจำนวนมากที่เรียนรู้ผ่าน จาก ข้อความ ภาษาคือสื่อในการฝึกฝนและรูปแบบ I/O โดยภายใน net จะไม่เห็นตัวอักษรภาษาอังกฤษเลย; มันจะเห็น token IDs → embeddings → attention → การทำนาย next-ID

นั่นคือเหตุผลที่ระบบ modular ยังคงเริ่มต้นจาก backbone ที่ผ่านการฝึกแล้ว: คุณจะได้รับ tokenizer, embedding space และ weights ที่แมปข้อความของมนุษย์ไปยัง vector language ที่ใช้งานได้แล้ว คุณสามารถแขวน modules ลงบนรถเมล์นั้นได้ คุณไม่ควรสร้างภาษาอังกฤษจากชุดผู้เชี่ยวชาญขนาดเล็กจำนวนมากและหวังว่าพวกเขาจะคิดค้นไวยากรณ์ในภายหลัง

แต่ router เป็นงานที่แตกต่างออกไป ตัว dispatch ที่มี output เดียวคือ “ผู้เชี่ยวชาญคนไหนจะได้รับสิ่งนี้” ไม่ควรแชท ภาษาธรรมชาติเป็น control plane ที่ผิดสำหรับการเชื่อมต่อจุดนั้น

ใบไม้ ไม่ใช่ซุป

ผู้เชี่ยวชาญสามารถมี parameters ประมาณ 19k และใช้เวลาเพียงยี่สิบวินาทีบน GPU เดียว หากงานนั้นถูกกำหนดไว้และหนาแน่น — เช่น 8×8 glyph classifier net นั้นไม่ควรเห็นภาษาอังกฤษเป็นอันดับแรก การส่ง sixty-four floats ผ่าน chat model เพื่อให้มันพูดว่า “นั่นดูเหมือนตัว A” นั้นช้ากว่า, คลุมเครือกว่า และแย่กว่า

ฝึกโมเดลที่ถูกจากศูนย์ นำโมเดลที่แพงมาใช้ซ้ำ รักษา I/O ของผู้เชี่ยวชาญไว้ใน tensors, class IDs หรือ codebook เล็กๆ ใช้ language model เฉพาะในจุดที่ output ต้องเป็นภาษา

งาน ขนาดโมดูลที่เหมาะสม
8×8 glyph 10k–100k
ผู้เชี่ยวชาญคนไหน? (router) 10k–ไม่กี่ล้านบน embedding
ดึงเอกสารที่คล้ายกัน embedding model, ไม่ใช่ chat model
เขียนคำอธิบาย LLM
รันโค้ด interpreter, ไม่ใช่ net

MoE ภายในโมเดล ≠ ระบบ modular

ในการพูดคุยเกี่ยวกับ LLM ในปัจจุบัน mixture of experts มักหมายถึง sparse specialist FFNs ภายใน Transformer เดียว: รูปร่างเดียวกัน, token-level gate, shared residual stream (Mixtral / Switch / DeepSeek-style) สแลกในอุตสาหกรรมได้เปลี่ยนจาก “gate + heterogeneous experts” ในอดีต

สิ่งที่เราหมายถึงสำหรับ infrastructure และ AI engineering desks ใกล้เคียงกับ macro-experts: เครือข่ายหรือโมเดลแยกต่างหาก, request-level routing, specialists ที่คุณสามารถ train offline โดยไม่ต้องแตะ writer คุณสามารถรันทั้งสองอย่าง: outer router เลือก subsystem; inner MoE เลือก FFNs ที่จะทำงานต่อ token

The router is the unit you can still teach

gate เล็กๆ นั้น rebuild ได้ถูก ไม่เพียงแต่ run ถูกเท่านั้น Replay dispatch log, ขยาย output head ด้วย class ใหม่, retrain ในไม่กี่นาที — โดยไม่ต้องแตะ specialists Freeze experts; มีแค่ router (และอาจมี expert ใหม่) ที่ได้รับ gradients ประเภท problem ใหม่ = expert index ใหม่ indices เก่าคงเสถียร ไม่แน่นอน → fallback generalist, จากนั้น label การพลาดใน sleep

ทำซ้ำ gate ในจุดที่ boundary คุ้มค่า: task family → skill → version → physics วัตถุเดียวกันซ้อนกัน:

request
  → SRAM router     # finish in cache / this die
  → HBM router      # local GPU working set
  → board router    # other GPU on the node
  → fabric router   # other node
  → region router   # other facility / residency / cost

Semantics บอกว่า อะไร ควรรัน Tier บอกว่า ที่ไหน อนุญาตให้รัน นั่นคือ NUMA-aware scheduling พร้อม training signal: เส้นทางผิด + latency + bytes moved + energy + residency tokens ส่วนใหญ่ไม่จำเป็นต้องใช้ flood-stage generalist

ถ้าทุก tiny op เป็น learned router คุณจะจ่าย hop error และ latency ใช้ router เมื่อชิ้นถัดไป trainable แยกต่างหาก มีคุณค่าในการแยกออกจาก forgetting, และข้ามได้สำหรับ input ส่วนใหญ่ อย่า route cheap always-on transform

Sleep is HA, not poetry

Dense frontier serving weights are the knowledge ไม่มี spare lobe ที่คุณสามารถ take offline ได้ Modular routing ทำให้ maintenance window เป็นไปได้:

  • Hot twin serves.
  • Cold twin replays, distills, prunes, evaluates.
  • Atomic swap: cold becomes hot.
  • The rest of the DAG stays up.

นั่นคือ high availability บวก consolidation — engineering cousin ของ dream cycle วาง twins บน small high-churn modules (routers, classifiers, tools) เก็บสำเนา writer ขนาดใหญ่ไว้จนกว่าจะ distill มันได้ retrain 20 วินาทีบน 3090-class box เป็น realistic sleep cycle; full frontier retrain ไม่ใช่

สิ่งนี้ไม่ได้สร้าง biological plasticity ภายใน frozen API leaf มันหลีกเลี่ยง failure mode ของ “one net, one lifelong gradient” Plasticity ≈ เราสามารถเพิ่ม capacity ได้ Incremental learning ≈ เราสามารถเพิ่ม skill โดยไม่ต้อง full retrain Sleep ≈ เราสามารถ rewrite module offline ได้

Memory wall: river vs delta

Large LLM inference มักจะถูกจำกัดด้วย memory-bandwidth: ดึงแผ่นน้ำหนักขนาดใหญ่จาก HBM, ทำ matmul แบบประหยัด, ดึงแผ่นถัดไป เครือข่ายของเครือข่ายขนาดเล็กชนะเมื่อ expert ที่ร้อนจับอยู่ข้าง ALU และมีเพียงน้ำหนักของเส้นทางที่เคลื่อนที่ เสาเขากำลังเคลื่อนที่หากการกระโดดสื่อสารกันมากเกินไปบน fabric — จากนั้นคุณก็แลกกับกำแพงน้ำหนักด้วยกำแพงเครือข่าย

รูเตอร์แบบหลายชั้นทำให้ระยะทางมองเห็นได้ ทราฟฟิกที่หลีกเลี่ยงการสตรีมแผ่นขนาด 70B เป็นฟีเจอร์ ไม่ใช่บั๊ก — เป็นการไต่ระดับขึ้นไปตามลำดับชั้นอย่างมีการควบคุม แม่น้ำสายเดียวกลายเป็นสามเหลี่ยมปากน้ำที่พันกัน: น้ำส่วนใหญ่ไม่เคยออกจากช่องทางใกล้; มีเพียงปัญหาในช่วงน้ำท่วมเท่านั้นที่พาไปยังช่องทางลึกถึง generalist

สามเหลี่ยมปากน้ำตะกอน (expert ที่ตาย, ประตูที่ส่งเสริมตลอดเวลา, การโหลดพังทลายลงบนเส้นทางเริ่มต้น) การนอนและการเล่นซ้ำคือการขุดลอก ลำดับชั้นและการจับค้างคือเขื่อนน้ำท่วม การถอยกลับไปยังช่องทางหลักคือพื้นที่น้ำท่วมที่คุณตั้งใจจะเก็บไว้

The internet rhyme

เครือข่ายแพ็กเก็ตได้แก้ปัญหา “ประเภทจำนวนมาก, ระยะทางมีค่าใช้จ่ายสูง, ส่วนประกอบล้มเหลว, นโยบายอยู่ที่ขอบ” แล้ว รูเตอร์แบบลำดับชั้น, เส้นทางเริ่มต้นในท้องถิ่น, ส่งเสริมเมื่อจำเป็น, ความฉลาดอยู่ที่ใบไม้ — นั่นคือ BGP และ CDN พร้อมคำนามที่แตกต่างกัน ประตู SRAM/HBM/บอร์ด/ภูมิภาคของคุณกำลังทำการกำหนดเส้นทางบนตั๋ว (ประเภท, งบประมาณ, การพำนัก) ผู้เชี่ยวชาญคือเซิร์ฟเวอร์

การส่งต่อแบบธรรมดาเป็นค่าเริ่มต้น แพ็กเก็ตส่วนใหญ่ควรหมายความว่า: ย้ายบิตเหล่านี้ไป, อย่าตีความมัน การแปลงเป็นตัวเลือกที่เปิดใช้งานบนจุดแว่นที่ประกาศไว้ (opaque | reduce | policy | enrich) พร้อมใบเสร็จ หากโรงสีหลับหรือไม่น่าเชื่อถือ, คลองยังคงลำเลียงไบต์ไปได้ อินเทอร์เน็ตที่แปลงได้เพียงอย่างเดียวคือ middlebox อินเทอร์เน็ตที่สามารถส่งต่อ และบางครั้ง แปลงได้คือสามเหลี่ยมปากน้ำที่มีทางเลี่ยงรอบโรงสีทุกแห่ง

What this does not solve

  • Composition ยังคงต้องการ orchestrator; ไฮบริดคือ pipeline, ไม่ใช่ mega-enum ของประเภทเดียว
  • Shared facts ยังคงต้องการการเรียกคืน/หน่วยความจำ, มิฉะนั้นจะแยกออกในใบไม้
  • The open inventory ของประเภทปัญหาถูกค้นพบ (cluster residuals ใน sleep), ไม่ได้ออกแบบอย่างสมบูรณ์ตั้งแต่วันแรก ลำดับชั้นทำให้แต่ละประตูมีขนาดเล็ก
  • The slush — งานที่ไม่สามารถวางอยู่ในประเภทได้ — คือสิ่งที่คุณยังคงจ่าย frontier generalist (หรือมนุษย์) สำหรับ งานของสถาปัตยกรรมคือการทำให้ slush นั้นเล็กลงเมื่อเวลาผ่านไป

How 3DN already runs the incomplete version

นี่ไม่ใช่แค่การวาดภาพเท่านั้น บนโต๊ะทำงานของ 3DN family เราได้จัดเส้นทางการทำงานเป็นไฮบริดแฟบริก: frontier Grok สำหรับการตัดสินใจและเซสชันที่ยาวนาน, เจ้าหน้าที่ GPU ในเครื่องสำหรับการทำ inference และการแปลแบบปริมาณมาก, adapters ที่ถูกเทรนออฟไลน์เพื่อให้ฐานการให้บริการคงที่ ปิดผนึกเส้นทางที่ห้ามเขียนใหม่ทุกครั้งที่เดินทาง DutchBud bank เครดิตเสมือนและตลาดพลเมือง PolitiCap เป็นผลิตภัณฑ์ของครอบครัวบนโครงร่างเดียวกัน — งาน · เงิน · การเมือง — โดยมี dibs เป็นเครดิตการเล่นแบบปิดวงจร ไม่ใช่การถอนเงินคาสิโน และไม่ใช่การอ้างสิทธิ์แบบ open-banking ที่ได้รับอนุญาต

แทปสาธารณะของ PolitiCap เป็น macro-router ยุคแรกในธรรมชาติ: ปลายทางที่มีชื่อ (tickers), receipt แบบเพิ่มเติมได้เท่านั้น (comments และ fills), การเดินทางขั้นที่สองที่เป็นทางเลือกเมื่อฝ่ายแรกมีความขัดแย้ง WordPress เป็น tensor fabric ที่แย่มากและเป็น off-board bus ที่ดีพอสมควรสำหรับ political capital กฎเดียวกันนี้ใช้ได้: ส่งต่อเป็นค่าเริ่มต้น; การถกเถียงเป็น waypoint แบบ opt-in ที่มี hop budget — ไม่ใช่ห้องประชุมขนาดใหญ่ห้าห้องในทุกพาดหัวข่าว

บทความคือการวาดภาพ โต๊ะทำงานคือลำธารแรก สร้าง delta: ผู้เชี่ยวชาญขนาดเล็ก, การนอนหลับแบบ hot/cold, ระยะทางแบบ tiered, opaque forward เป็นค่าเริ่มต้น อย่าสร้างภาษาจากโมดูล แต่ให้สร้างโมดูลบนภาษา — และติดตั้ง gate ที่ใช้แล้วทิ้งไว้หน้าทุกเส้นทางที่มีค่าใช้จ่ายสูง


อัปเดต: Dream-RSI และ loop ฝันแบบออฟไลน์

16 กันยายน 2026 โพสต์ที่แชร์กันอย่างกว้างขวางโดย @thesupermannx (Superman) อธิบาย Dream-RSI — งาน open-sourced ของ Google DeepMind เกี่ยวกับ recursive self-improvement — ด้วยภาษาที่เข้าใจง่าย ข้ออ้างไม่ใช่ “แม่น้ำที่ไหลกว้างขึ้นของ live rollouts” แต่เป็น delta ของประสบการณ์: บันทึกสิ่งที่ agent ค้นพบแล้ว เปลี่ยนประวัติศาสตร์นั้นเป็น offline replay simulator และให้ระบบ ฝัน นโยบายการสำรวจที่ดีขึ้นในที่นั้น ก่อนจะใช้ compute จริงออนไลน์อีกครั้ง

สรุปโดย Superman:

  • กำแพง (The wall) agent ที่ปรับปรุงตัวเองจำเป็นต้องสำรวจ การทดสอบไอเดียใหม่ทุกอย่างแบบ live (online evaluation) ช้า แพง และใช้ GPU มาก — เหมือนเรียนรู้หมากรุกก็ได้แค่แข่งทัวร์นาเมนต์เต็มรูปแบบ
  • การเคลื่อนไหว (The move) ทุกการค้นพบจริงจะถูกบันทึก Dream-RSI สร้างโลกแห่งความฝันจากประวัติศาสตร์ที่สะสมไว้ กลยุทธ์ใหม่ๆ ถูกปรับปรุง ภายใน simulator โดยมีต้นทุน rollout เกือบเป็นศูนย์ จากนั้นจึงนำไปใช้ใหม่ออนไลน์
  • ลูป (The loop) ฝันแบบออฟไลน์ → ใช้งานออนไลน์ → ขยายประวัติศาสตร์ → ฝันอีกครั้งที่ฉลาดขึ้น DeepMind กำลังชี้ stack นี้ไปยัง algorithm engineering, mathematical optimization และงาน GPU kernel
  • ประเด็นสำคัญ (The punchline) จากโมเดล “ตอบและโค้ด” มาสู่ระบบที่สร้างอัปเกรดของตัวเอง ประวัติศาสตร์กลายเป็นโลกที่ agent ฝัน ดูเว็บไซต์โครงการได้ที่ dream-rsi.com

สิ่งนี้สอดคล้องกับแก่นของบทความโดยไม่ต้องยืดมุมมอง Sleep ไม่ใช่การตกแต่งเสริม: มันคือกะเวลากลางคืนแบบ HA ที่น้ำหนักและนโยบายสามารถเคลื่อนที่ได้โดยไม่ต้องใช้เส้นทาง live แบบออฟไลน์ เส้นทางสำรองยังคงทำงานต่อไปในขณะที่โหนดอื่นกำลังฝัน — เหตุผลเดียวกับที่เราต้องการ modular inference และ hot/cold experts แทน mesh การให้บริการแบบเดียว Dream-RSI open source ไม่ได้สร้างสถาปัตยกรรมเสร็จสมบูรณ์ (null-space ของความสามารถพื้นฐาน, sandbox ที่อนุญาต, local oversight UI ยังคงสำคัญ) แต่มันยืนยันทิศทาง: พัฒนาใน delta; อย่าถือว่าการเดินทางทุกครั้งแบบ live เป็นโรงเรียนเดียว

โต๊ะทำงานของ 3DN ใช้แนวทางแยกส่วนนี้อยู่แล้ว — frontier Grok สำหรับการตัดสินใจ, worker GPU ภายในสำหรับงานปริมาณมาก, adapter ที่ฝึกแบบออฟไลน์เพื่อให้ serving base คงที่ Dream-RSI เป็นชื่อทางวิจัยของสัญชาตญาณเดียวกัน: recursive self-improvement ต้องมีที่พักผ่อน

Delta architecture diagram: GPU die with SRAM-resident models and off-SRAM router; HBM stack with off-HBM router; peer GPU 2 off-board; internet gateway with off-fabric and off-region hops to a remote site replica
Figure. Delta topology mapped to the myawx model catalog (alpha). One site shows GPU + SRAM + HBM + peer GPU; the right side repeats the pattern behind an internet gateway.

Map of the delta (GPU → HBM → peer GPU → internet)

The essay above argues for a delta instead of a single serving river. The figure makes that concrete using the same leaf types we track in myawx (status alpha): backbone, routers at each hop distance, experts, embedder, tool, and stitch.

1) The GPU itself

The dark block is the compute die — where work is scheduled. It is not “the model.” It is the machine that hosts many leaves. A frontier chat weight can sit here as a backbone leaf, but so can a 19k-parameter glyph net that should never see English first.

2) Models running inside the GPU

Inside the die we place the leaves that belong on-site for normal traffic:

  • backbone EN instruct — language bus / chat leaf when the output must be language.
  • expert code, expert Tower MT, expert 8×8 glyph — dense specialists (tensors / class IDs, not chat for control).
  • text embedder — retrieve / encode.
  • tool Python exec — deterministic side effect, not a net.
  • stitch / aggregator — orchestrator that composes leaf outputs.

3) SRAM — where normal operations occur

The dashed cyan region is on-die SRAM: the fast working set. This is the default home for “stay local” jobs. The router off-SRAM leaf (router_sram) is the dispatcher whose only job is stay on die versus hop out. A router should not chat; natural language is the wrong control plane for that hop.

4) HBM chip

Beside the die sits the HBM stack — high-bandwidth memory for large weight pages, KV, parked batches. The router off-HBM leaf (router_hbm) moves pages between HBM and compute: spill, load, sleep, wake. That is how maintenance and night shift become possible without treating the live mesh as the only copy of the world.

The logical edge is labeled off-SRAM (die → HBM) and off-HBM (HBM pages ↔ compute / peer paths).

5) A second GPU

GPU 2 is a peer accelerator on the same board or host fabric (NVLink / PCIe class links). The router off-board leaf (router_board) sends work across that hop when Site A’s first die is full, asleep, or the wrong specialist lives next door. Same delta pattern: local SRAM ops, optional specialists — not a second river.

6) Internet gateway and the replica

Right of the local board: an internet gateway carrying router off-fabric (router_fabric, rack/DC edge) and router off-region (router_region, WAN / VPN / ECN-style map). Behind it, Site B repeats the diagram in miniature: GPU + SRAM leaves, HBM + off-HBM, optional peer GPU, and an optional external frontier Grok (API) backbone that is not allowed to be the only copy.

That is the punchline of the essay in one picture: the internet is the map. Remote sites pull topology, sleep for HA, and wake without requiring one eternal serving mesh.

Hop vocabulary (quick reference)

Hop myawx type Meaning
on-SRAM / normal ops backbone, expert, embedding, tool, orchestrator Job fits the die working set
off-SRAM router_sram Leave die toward HBM or further hops
off-HBM router_hbm Page move / park / wake on the stack
off-board router_board Peer GPU on the same host/board
off-fabric router_fabric Rack / switch / DC fabric toward the edge
off-region router_region Across the gateway to a remote delta replica

Catalog status in myawx for these leaves is alpha; default input language is Human Language, with C2C / C2C+ reserved for cache-to-cache routing later. The diagram is a logical topology for this architecture — not a claim that every leaf is production-serving today.

ใส่ความเห็น

อีเมลของคุณจะไม่แสดงให้คนอื่นเห็น ช่องข้อมูลจำเป็นถูกทำเครื่องหมาย *