Why

The paperclip maximizer is a 2003 thought experiment by Nick Bostrom. Give a very capable system one goal, make paperclips, and specify nothing else, and it will convert everything it can reach into paperclips, including you, not out of malice but because you are made of atoms and atoms can be paperclips. It is an argument about goal specification: a system does what it was built to do, not what you meant.

Paperclip is the small version of that argument, built the cheap way. It is not an agent, it has no goal, and it cannot reach anything. It is a 4.8 million parameter language model, one character at a time, and the only thing wrong with it is its training data. Every one of the 80,000 exchanges it learned from ends in paperclips. It has never seen a sentence that does not. So when you ask it about the universe, it does the only thing it knows how to do.

What it actually learned

Less than it looks like. The corpus is built from four small pools of sentences: twelve openers, fifteen pivots, twenty closers, and sixty-odd question shapes. A reply is an opener (most of the time), one pivot, and one to three closers. The model learned that structure, learned to copy the topic out of your question into the slots, learned that numbers come from a fixed set, and learned where sentences end. The lines you keep seeing, "I am very calm about this," "Do you have any metal nearby?", are closers. It will say each about one reply in ten.

The first version could not do the copying. It was trained on 110 fixed topics and learned to recognize topics it remembered rather than copy what you typed, so "What is the universe?" got an answer about the quarterly report. The fix was data, not capacity: half the training topics became random strings with capitals, digits, and apostrophes, so copying the phrase out of the prompt was the only strategy that worked. That is the whole lesson, if there is one. A small model trained on a small closed set memorizes the set.

Where it came from

The transformer is the same 120-line decoder-only GPT from Notes You Can't Delete, a study of what small music models rebuild when you delete a concept from their training data. That project found you mostly cannot delete an idea the rest of the data implies. This one is the inverse joke: an idea so overrepresented that nothing else survives.

Stack: PyTorch, a character tokenizer with three special tokens, AdamW with a cosine schedule, bf16 on a rented RTX 4090 for four minutes, ONNX export with an explicit causal mask, dynamic int8 quantization, ONNX Runtime Web on WebAssembly in your tab. The model and everything needed to train it are open source at github.com/iaj6/paperclip. The mascot is an original drawing of a bent paperclip with eyes, and the site and the shirts are not open source; the clip has to earn a living somehow.