Hiroki Butterfield

Building small language models (SLM) part 1

Or small models in general.

In a world of constantly larger models held by private institutions, I thought it might be better to focus primarily on what can be done with constraints. It’s absolutely ridiculous to me to think that we need 100s of millions of dollars to get “intelligence”. Constraints are a good thing. They often:

  1. breed innovation
  2. make disconnected/offline use practical.

Plus, if you want to do anything in the realm of physical intelligence, you’re going to be operating in a hardware constrained environment.

I personally want something that I can

  1. train on my own computer
  2. run on my own computer
  3. find valuable (and possibly share with others)

In general, I think development under constraints can yield interesting work. Plus, thanks to the work of LLM providers today, we can get further than we might have starting from nothing.

What is an SLM?

Wikipedia says these typically have fewer than 40 billion parameters, but there’s no universally agreed cutoff. That number alone doesn’t tell you what can be trained or hosted on consumer hardware. Small models can be trained from scratch, or trained using a larger model as a teacher through distillation (SmolLM). Quantization is another way to reduce the memory needed to run a model, without changing its parameter count. The real challenge or question for me is what can be trained on smaller compute. Running a model locally, fine-tuning it, and pretraining it from scratch are very different resource commitments.

For an overview of SLMs, model compression, and running models on consumer devices, see John Johnson’s Small Language Models (SLM): A Comprehensive Overview (Hugging Face community article).

General models today

At a high level, a deployed AI system combines the model with surrounding software, sometimes called a harness, that handles things like prompts, tools, and execution. For this discussion, I’m primarily focusing on just the model.

Model development typically goes through the following phases:

Data acquisition
    ↓
Data preparation
    ↓
Tokenizer development
    ↓
Model design and training setup
    ↓
Pretraining
    ↓
Mid-training (Optional capability adaptation)
    ↓
Post-training

Mid-training can include things like extending context length or developing reasoning capabilities; evaluation happens throughout training, rather than being a separate stage (SmolLM3 training recipe). Not every project uses all these stages.

Small models can be constrained across a number of dimensions.

Data acquisition

Firstly, there’s the amount of data in question. If you’re talking about a smaller machine, you cannot be reasonably looking at the entire internet. Not to mention the absolute metric ton of expensive expert data and training environments that labs like Anthropic and OpenAI are buying. OpenAI reportedly budgeted around $1 billion for data in 2025, including human experts and reinforcement-learning environments (The Information, September 2025). Anthropic reportedly discussed spending over $1 billion on reinforcement-learning environments over the following year (Epoch AI, citing The Information).

These budgets cover more than pretraining data: reinforcement-learning environments are used in later training stages too.

While it boggles the mind that paying for all this expert data has led to material progress, I personally would expect an intelligent system to be able to learn autonomously and not require purchased data, though that’s not a real blocker to a system being useful.

For our purposes, we’ll just use whatever is simply available publicly.

Data preparation

Data preparation is the dirty plumbing work required to get data in a place where it can actually be used for pre-training. This is everything like extraction, normalization, filtering out spam and “bad stuff”, deduplication, splitting for training etc. Not much to say here. This grunt work likely has to happen no matter what. I do wonder if labs have gotten to a point where they just use their own models to do this work.

Tokenizer development

Tokenizer development is how you get the vocabulary and rules that convert text into token IDs and back. BPE (byte-pair encoding) with a fixed vocabulary is common in language models. I don’t think re-inventing the wheel here is necessary. Laya’s multilingual version, for example, supports 100+ languages which evidently covers their use case overall by re-using mmBERT’s existing tokenizer and pretrained model (Laya model card). Laya is an encoder-based decision model, so it’s a tokenizer example here rather than the same generative architecture we’ll discuss below.

However, Sarvam developed a custom tokenizer for its Sarvam-1 language model to represent Indian languages more efficiently. Its vocabulary contains 68,096 tokens, including 4,096 reserved for future expansion, and Sarvam reports an average of 1.4–2.1 tokens per word across its 10 supported Indian languages (Sarvam-1: tokenizer details).

Given we’ll mostly be operating in English, a pre-existing tokenizer is probably #justfine if we’re training from scratch - no sense in adding more difficulty than necessary. I’d still check its vocabulary size and how efficiently it represents our data: a large vocabulary can eat up a small model’s parameter budget. Chinese models also use BPE: Qwen3-8B has a base vocabulary of 151,643 tokens, while DeepSeek-V3 has 128,000, with additional special tokens handled separately (Qwen3 tokenizer, DeepSeek-V3 tokenizer).

Counts below are unique token IDs in the published tokenizer files:

Model Tokenizer Vocabulary size Counting detail / source
Laya — multilingual mmBERT BPE 256,000 Includes special tokens. Tokenizer file
Sarvam 105B BPE 262,144 Includes special tokens. Tokenizer file
DeepSeek-V4.1-Flash BPE 129,280 128,000 base entries plus 1,280 additional tokens. Tokenizer file
Qwen3.8-Flash-Next BPE 248,077 248,044 base entries plus 33 additional tokens. Tokenizer file

Vocabularies do seem to be getting bigger over time as well. Sarvam went from about 68k to 262k entries, and Qwen from about 152k to 248k (Sarvam-1, Qwen2.5 tokenizer).

Mistral says its newer Tekken tokenizer does this more efficiently (Mistral NeMo). Bigger isn’t free though: those extra entries need bigger embedding and output tables. And newer doesn’t necessarily mean fewer tokens either. Anthropic says its newer tokenizer produces roughly 30% more tokens for the same text, depending on the content (Anthropic docs).

So there seems to be a trend here, but I wouldn’t take vocabulary size alone as a measure of how good a tokenizer is.

It’s unclear what this trend implies. A fixed byte-level tokenizer can represent more languages without a bigger vocabulary, but it may need more tokens to do so. A bigger vocabulary can make that more efficient. I also wonder whether finer token granularity helps in some domains, but Anthropic’s announcement only establishes the change in token counts, not a performance benefit.

Model Design and Training Setup

Model design is interesting within the context of smaller models because choices like vocabulary size can take up a much bigger share of the parameter budget, even if you’re using a vanilla GPT-2 decoder-only Transformer.

Let’s take tokenization for example. If you have a vocabulary of 200k tokens, which is about the size for many of these SOTA models, and the embedding table (the token-vector map) is 512 dimensions, you’re looking at 200,000*512=102,400,000 or 102 million parameters for just the token embedding table!

Here are some actual embedding table sizes. These counts use the model configuration’s vocabulary size, which can include reserved or padded rows beyond the tokenizer’s entries. The Qwen model here is also different from the one in the tokenizer table above. These are input token-embedding tables, not total model parameters or combined input/output tables.

Model Embedding-table rows Embedding width Embedding-table parameters
SmolLM2-135M 49,152 576 28 million
SmolLM2-1.7B 49,152 2,048 101 million
Qwen3.8-27B 248,320 5,120 1.27 billion
Sarvam 105B 262,144 4,096 1.07 billion
DeepSeek-V4.1-Flash 129,280 5,120 662 million

Let’s take nanochat for example (Karpathy’s successor to nanoGPT). Assuming you’re using nanochat’s d12 configuration as our baseline (model configuration, training setup) you have a number of ways to configure your model:

Hyperparameter Value Meaning
Depth (Transformer blocks) 12 Number of Transformer blocks.
Model width 768 Size of the main token vectors. The dimension stays the same across blocks; their weights are separate.
Attention heads 6 Number of query heads in each attention layer.
Dimensions per head 128 Size of each head’s vectors.
Key/value heads 6 Number of key/value head pairs in each attention layer.
Feed-forward hidden width 3,072 Size of the expanded layer inside each block’s feed-forward network.
Context length (tokens) 2,048 Maximum number of tokens in a training sequence.
Default tokenizer vocabulary size 32,768 Number of different token IDs in the default tokenizer.

In this version of nanochat, that comes to roughly 286M parameters. The configuration table doesn’t tell the whole story: six extra value-embedding tables account for about 151M of those parameters, and there are also value-embedding gates1.

And then you have all sorts of training considerations:

Consideration Description
Starting point Random weights, continued pretraining, or fine-tuning an existing model?
Training objective What counts as a correct prediction, and which tokens contribute to the loss?
Data mix and ordering How much general text, code, domain data, and repeated data does it see?
Sequence packing How do documents become batches of token sequences?
Batch size How many sequences are processed together? Some setups count tokens instead. The effective batch per update can include several smaller batches through gradient accumulation.
Gradient accumulation How do you reach a larger batch when memory is limited?
Optimizer What rule turns gradients into weight changes?
Learning rates and schedule How large are updates, and how does that change over the run?
Regularization and stability How do you discourage overfitting and keep updates well behaved?
Numerical precision Which calculations and stored values use FP32, BF16, or another format?
Memory and execution What fits, and how efficiently can the device process it?
Training budget How many tokens or updates will you train for?
Evaluation How will you tell whether training is improving the model?
Checkpointing and reproducibility How can you resume or compare runs?

What’s next

I’m not going to go into training in this article but the point to consider is that large model providers are tuning many hyperparameters and have both enormous budgets and powerful computers to do so. They operate on a completely different optimization space than someone trying to build in a more constrained environment. I find DeepSeek’s work on inference and training efficiency interesting in that context. I suspect hardware constraints can push teams toward useful optimizations, though that isn’t enough to establish which changes were driven by geopolitics.

Footnotes

  1. I don’t actually know what these do ↩