Logbook: the experiment ended before the download did
2026·09·27 · 2 min read

Photo: Nemuel Sereti · Pexels
I wanted to run Ling-3.0-flash-VL locally — 124B parameters under an MIT license sounds like something you can just pull. Reading the launch matrix before starting the download saved me 250 GB and an afternoon. That is the result.
The plan
Pull Ling-3.0-flash-VL — published by inclusionAI on September 4, 2026 under an MIT license — serve it with vLLM or SGLang, and find out what the real VRAM requirement is versus what the model card promises.
Step 1: read the card before touching git lfs
This is the step I usually skip. The card's own numbers:
- 124B total parameters, 5.5B activated per token (sparse MoE).
- Image and video input, context window up to 256K tokens.
- License: MIT. Genuinely MIT, not a community license with a seat cap.
And then the serving section, which is where the afternoon ended. The recommended SGLang recipe for the full 256K context (via YaRN) is four 141GB-class GPUs — H20-3e or H200 — or a four-GPU Blackwell node (B300 / GB300). There is an FP8 variant published four days later that lowers the bar, but not to anything that exists in this room.
Step 2: check what "download the weights" means
The repository's file metadata puts the BF16 checkpoint at roughly 250 GB across 64 shards. On this connection that is most of a day, to produce a directory I cannot load.
The finding
The experiment failed, and the failure is the useful part, so here it is plainly: an MIT license tells you what you are allowed to do, not what you are able to do. "Open weight" in 2026 means you can inspect, fine-tune and resell the model without asking anyone. It does not mean the model fits on your hardware, and for anything at this scale it does not.
The realistic path for a model like this, for someone without a GPU cluster, is still an inference provider's API — the difference from a closed model being that there is more than one provider serving the same weights, so switching is a base URL and not a rewrite. That is a real benefit. It is just not the one the word "local" implies.
What I would need for the actual experiment
Access to a four-GPU H200-class node for an afternoon. If that materialises, this entry gets a part two with the numbers the card promises measured against the numbers the machine reports. Until then I am not going to pretend I ran it.