Tag: Llm

Jev-Like Models

TypeSafe’s Jev has sparked a lot of community activity. I tried to collect a few and list them here:

The blossom of these Jev-like models shows there are huge technical challenges in matching Jev’s bench results, though Jev still leads on accuracy and speed. But Jev is not irreplaceable. For example, you can trade accuracy for speed with Reflex.

The bench has a full list of more Jev-like models you can check out.

C2C Links Models Through Their KV-Caches

Cache-to-Cache (C2C) enables large language models to communicate directly through their KV-Caches, bypassing text generation. By projecting and fusing KV-Caches between models, C2C achieves 8.5-10.5% higher accuracy than individual models and 3.0-5.0% better performance than text-based communication, with a 2.0x speedup in latency.

This is an interesting experiment. The C2C approach removes the intermediate tokens and goes straight for “thought projection.” After Model A computes, it doesn’t generate any text at all. The system uses a lightweight neural network (the Neural Fuser) to splice and fuse Model A’s attention memory (KV-Cache) directly into Model B’s internal KV-Cache, through high-dimensional spatial rotation and alignment. The challenge is that different models have varying numbers of layers and structures. In the paper, the model dynamically senses on its own which key layers absorb the highest gains from external caches, and which layers should stay independent in thought, with millisecond-level adaptive balancing.

This is quite similar to Mostik. C2C states clearly it’s passing KV-Cache: it trains a projector plus a cache fuser plus a gate, and fuses the source KV into the receiver’s KV before decoding. Mostik talks about a more general “hidden state,” trains a bridge to map the sender’s latent space to the receiver’s, then lets the receiver keep generating. So far, C2C seems more practical and has more details than Mostik.

Open Takes on Jev: SemIf and Laya

OpenJev has been renamed SemIf.

It’s wise to separate it from Jev’s marketing. It’s an independent research project that just provides the same API. The underlying architecture might be completely different, since TypeSafe never disclosed Jev’s own model or training.

SemIf uses Qwen and MiniCPM5 as its core models. Instead of an autoregressive decoder, it does one forward pass and reads probabilities over filtered logits.

Another project, Laya, is also worth a look. It claims to outperform Jev by a tiny margin. The model is small, 421M params. Laya evaluates typed questions (choice, score, noul) over any state, text, email, ticket, or a JSON document, in a single forward pass too.

I ran it myself on an M4 Pro. Without preloading, it took 66 seconds to classify a code-change type: given a PR diff, check whether it’s a feature change, a docs change, a test change, an IaaS change, a version bump, and so on. I gave it a version-bump-only diff, and it scored version bump at 0.7477, with everything else below 0.55. It classifies well. It already looks usable.

Reading GuppyLM

GuppyLM is a small language model that talks like a fish. The repository is small enough for a beginner to read from end to end. It shows the path from data and tokenization to training and inference without hiding the pieces inside a large framework.

It is slightly more complex than Karpathy’s microgpt. That is useful. GuppyLM is not an implement-everything-from-scratch demo. It uses PyTorch, separates the model, dataset, training loop, and inference code, and looks closer to the code used to train and serve models today.

The model is still simple: 8.7 million parameters, six Transformer layers, six attention heads, a 4,096-token vocabulary, and a 128-token context window. The training code includes AdamW, learning-rate warmup and cosine decay, mixed precision, gradient clipping, evaluation, and checkpoints. The inference code loads the tokenizer and checkpoint, then generates tokens with temperature and top-k sampling.

The model trains on 60,000 synthetic conversations. The project says training takes about five minutes on one GPU. Its quantized ONNX export is about 10 MB and runs in a browser. This makes the whole loop fast enough to inspect, change, train, and test instead of only reading about it.

Testing PrismML Bonsai 2

PrismML released Bonsai 2 27B this week. It is built on Qwen3.8 27B, but compressed to ternary weights: each weight is +1, 0, or -1 instead of 16 bits. That drops the size from about 56 GB to 5.9 GB. On PrismML’s benchmark suite, it keeps about 98% of the original model’s score. It runs on a Mac (Metal), on Linux or Windows (CUDA, Vulkan, ROCm), or on CPU alone.

I tested it on my M4 Mac with 24 GB of RAM. I wired it to my Pi coding agent and gave it a code review task. It is pretty cool.

Bonsai 2 followed my codereview.md rules. It checked the change against the other files in the repo, the way I asked it to.

The full task took 30 minutes. Token speed on my M4 was about 10 tokens per second. Other reports show 40+ tokens per second on an M5.

Testing Jev: Three Playground Runs

I got early access to Jev after writing about it 2d ago. Here are three tests I ran in the playground, with the actual state, questions, and answers.

Is a Jedi sandwich a sandwich?

State:

{
  "food": "Jedi sandwich",
  "definition": "Put Luke Skywalker in between two slices of toast."
}

Question:

{
  "is_sandwich": {
    "type": "noul",
    "instructions": "Is `food` a sandwich?",
    "criteria": {
      "true": "A sandwich is a food dish where a filling, such as meat, cheese, vegetables, or spread, is placed between structural starch",
      "false": "The food has no bread enclosing a filling or uses only a single slice of bread, or uses a non-bread wrapper such as a tortilla, wafer, or cookie."
    }
  }
}

Answer:

[…723 words]

Jev: State In, Typed Decisions Out

Diogo Almeida posted on X about a new model, Jev, from a company called TypeSafe. TypeSafe calls it its first “System One” model, a term borrowed from Daniel Kahneman’s split between fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning. A normal LLM writes out its reasoning in text. Jev skips straight to a typed answer.

What it actually does

Ask a normal LLM “given this incident, what should we do?” and it answers in text:

[…959 words]

Mostik Links Models in Latent Space

Sasha Malysheva, Mostik’s CEO, announcing a new approach for model communications.

They connect a large model and a small model directly through their internal representations, instead of making them communicate with text.

In one experiment, they linked GLM-5.2 with Qwen-3.5. The result performed between the two models, at about 1/20 the cost of running the large model alone.

The idea is simple: let the small model handle most of the work, and use the large model only for the hardest reasoning.

If this works, future AI systems may not be one giant model, but networks of models that share internal state directly.

Pretraining

Pretraining is the first stage of training a language model. Its main task is to predict the next token. Take “The capital of France is ___.” A pretrained model reads “The capital of France is” and predicts the next token, “Paris.”

The model makes a prediction, compares it with the real token, calculates the error, and updates its weights.

To pretrain a model, you repeat this step over and over, on a huge pile of text, code, and other data:

Animated diagram of the pretraining loop: tokens feed a prediction, the prediction is compared to the real token to get a loss, the loss updates the model’s weights, and the loop repeats.

After pretraining, the model has picked up language, facts, code, and common reasoning patterns from that data.

Pretraining is only the first stage. Compare it with the two stages that usually follow:

StageGoalDataFeedback signal
PretrainingLearn general language patternsHuge, mixed text and codeNext-token prediction error
Fine-tuningLearn a specific task or formatSmall, curated examplesDifference from a labeled output
RLLearn a preferred behaviorThe model’s own outputsA reward score

Pretraining teaches a model what patterns exist in its data. Feed it a lot of code, and it gets better at code. Feed it many tool-call examples, and tool use can become a natural part of its output.

RMSNorm

Most modern language models, like Llama and Mistral, normalize activations with Root Mean Square Normalization (RMSNorm). Here is what it does, and a PyTorch module you can drop into a model.

The problem it solves

Inside a neural network, a layer’s output feeds the next layer as input. Those values can become too big or too small as they pass through layer after layer. A value that doubles at each of 40 layers is unusable by the end. Normalization rescales the values back to a steady range before they move on, so training stays stable.

[…388 words]

Train A BPE Tokenizer

I built train_tokenizer: a CLI that trains and evaluates a byte-level BPE tokenizer. It has 100 rows of sample data and a test file.

What a tokenizer does

A language model does not read text. It reads a list of numbers. A tokenizer turns text into that list, and turns the list back into text.

The simplest tokenizer splits text on spaces, into words, and gives each word a number. This breaks fast. Any word the model has not seen has no number. A model that never saw “photosynthesizing” cannot represent it.

[…1152 words]