Why Plugable (and Microsoft) Are Betting on Smaller AI Models
Product Owners | October 08, 2026
TL;DR
- On October 7, 2026, Microsoft launched the Surface RTX Spark Dev Box, a desktop machine built to run AI locally that starts at $5,999.99.
- Microsoft says GitHub Copilot can hand off up to 20% of its workloads to smaller models.
- Plugable's TBT5-AI16 pairs a Thunderbolt 5 enclosure with an NVIDIA RTX 5080 and connects to the Windows 11 laptop you already have, over a single cable.
- Bernie Thompson, Plugable's founder and CTO, calls it intelligence compression: "It's common to use a big model to create the skills that your small models will benefit from, over and over."
- Plugable's free VRAM Modeler shows how much GPU memory a model needs and whether it fits.
On October 7, 2026, Microsoft launched the Surface RTX Spark Dev Box. It's a desktop machine built to run AI locally, and it starts at $5,999.99. When a company Microsoft's size builds dedicated hardware for local AI, it's a clear sign of where a lot of AI work is heading: onto hardware you own.
We agree, and we've been shipping that idea since the first half of 2026. Our TBT5-AI16 pairs a Thunderbolt 5 enclosure with an NVIDIA RTX 5080 and connects to the Windows 11 laptop you already have, over a single cable.
It's a good day when a company Microsoft's size agrees with you. We'll try to stay humble about it.
You don't need the biggest model for every job
Microsoft's product page covers the aluminum chassis and the 1,000 air vents. The more interesting detail is about routing. Microsoft says GitHub Copilot can hand off up to 20% of its workloads to smaller models, and it pitches running AI locally as a way to cut per-token API costs.
In other words, a big chunk of everyday AI work can go to a smaller model. So why are smaller models suddenly good enough to do real work? Bernie Thompson, our founder and CTO, has been thinking about exactly that. He's been tracking how quickly open models are catching up to the big cloud models, and for coding and agent work the gap is closing fast.
Bernie on "intelligence compression"
Bernie recently shared a short post on LinkedIn about what he calls intelligence compression. It starts with the big models:
"Frontier Models distill much of humanity's accumulated knowledge, and compresses it into a few trillion parameters representing things, concepts, meta-concepts, etc."
Then the models get smaller, small enough to run on a desk:
"Smaller production and sub-models in a mixture-of-experts structure, further distill and quantize (drop some depth of detail) to compress that knowledge even more. Some of these run on local hardware."
"Quantize" is the step that decides whether a model fits on your GPU. Trim a little detail, and the same model fits in much less GPU memory. (Our guide to GPU memory for local LLMs explains the trade-offs.)
A smaller model knows less on its own. Bernie's point is that it doesn't have to work alone. A harness is the software around a model that lets it take several passes at a problem: plan, try something, check the result, and try again.
"By letting those smaller models think during a single turn, and also explore, execute, and test over many turns with a harness, they can extrapolate and discover nearby problem-specific and private knowledge and compress that, too, into the context."
That's where skills come in. A skill is your team's know-how written down ahead of time, as data, code and instructions the model can use. (If you want the mechanics of connecting AI to your data, our post on RAG and MCP walks through it.)
"At the company level, it's possible to create a large collection of skills which capture most of the institutional knowledge of a company, letting models answer questions, reason, and take action over the entire business. And because this knowledge comes pre-compressed, it sets small models up for success, and with less tokens used."
His last practical point is the one we'd underline twice:
"It's common to use a big model to create the skills that your small models will benefit from, over and over."
So you pay for the big model once, to write down what your company knows. The small model on your desk uses that work every day. That lines up with Microsoft's 20% routing figure.
What this changes about work
Plugable's CEO, Lynn Smurthwaite-Murphy, comes at the same idea from the business side. Her white paper, AI is a Change Agent, Not a Tool, argues that "context is the bridge between AI activity and AI value."
She also makes the case for local hardware without mentioning a single spec: "Much of a company's most valuable intelligence is hidden in the data public AI cannot see." Sometimes the answer is to bring the AI to where that data already lives.
Bernie explains how small models can work with what your company knows. Lynn explains why that changes how people work. Her paper covers strategy, team adoption and the people side of change, and it's worth reading in full: AI is a Change Agent, Not a Tool.
Where you can start
If you'd like to try this on your own desk:
- Size it first: Our free VRAM Modeler lets you pick a model, a quantization level and a context length, then shows how much GPU memory it needs and whether it fits.
- Hardware: The TBT5-AI16 adds 16GB of GPU memory to a Windows 11 laptop with Thunderbolt 5, Thunderbolt 4 or eGPU-capable USB4. For bigger jobs, there are 32GB and 96GB versions. See the full lineup.
- Models: It runs on Microsoft Foundry Local, part of the same Microsoft ecosystem behind this launch. Here's our introduction to Foundry Local.
- Software: The Enterprise Series includes Plugable Chat, a local app for asking questions of your documents and data in plain English. Version 1.0 is in final testing now.
- Honest limits: It's Windows 11 only (no macOS, Linux or ChromeOS) and needs an 11th-gen Intel or Zen 4 AMD processor or newer.
Questions about fitting it into your setup? Our North American team is at sales@plugable.com.
One last image
"Like a figure skater compressing their body to the center, the more we use AI to compress our knowledge, the faster our business is capable of spinning."
FAQs
What is a harness in AI?
A harness is the software around a model that lets it take several passes at a problem: plan, try something, check the result, and try again.
What is a skill in AI?
A skill is your team's know-how written down ahead of time, as data, code and instructions the model can use.
What does it mean to quantize an AI model?
"Quantize" is the step that decides whether a model fits on your GPU. Trim a little detail, and the same model fits in much less GPU memory.
How do I know if an AI model will fit on my GPU?
Our free VRAM Modeler lets you pick a model, a quantization level and a context length, then shows how much GPU memory it needs and whether it fits.
What computers work with the Plugable TBT5-AI16?
The TBT5-AI16 adds 16GB of GPU memory to a Windows 11 laptop with Thunderbolt 5, Thunderbolt 4 or eGPU-capable USB4. It's Windows 11 only (no macOS, Linux or ChromeOS) and needs an 11th-gen Intel or Zen 4 AMD processor or newer.
Related Articles
- What Is Stable Diffusion? Local AI Image Generation with Plugable TBT5-AI
- GPU Comparison for Machine Learning Workloads
- Local AI vs Cloud AI: Why Businesses Are Moving Toward Hybrid AI
- Bernie Thompson of Plugable Named to Channel Insider's 2026 AI Leaders in the Channel List
- GPU VRAM Requirements for Local LLMs | Plugable Guide
Loading Comments