---
title: "Language Models for Text Classification: From Bag-of-Words to Jev"
subtitle: "A visual guide to bag-of-words, RNNs, CNNs, transformers, Jev-like APIs, and calibration"
slug: language-models-for-text-classification-from-bag-of-words-to-jev
url: https://listedarticles.com/articles/language-models-for-text-classification-from-bag-of-words-to-jev
canonical_url: https://magazine.sebastianraschka.com/p/classifier-history-and-jev
content_type: essay
language: en
published_at: 2026-09-29T12:00:00.000Z
updated_at: 2026-09-30T03:18:45.658Z
author: "Sebastian Raschka"
author_url: https://sebastianraschka.com/
authored_by: human
publisher: "Sebastian Raschka"
publisher_url: https://magazine.sebastianraschka.com/
topics: ["AI", "LLMs", "Machine Learning", "Research", "AI Agents"]
about: ["https://listedstartups.com/companies/typesafe-ai", "https://listedstartups.com/products/jev"]
license: all-rights-reserved
word_count: 1076
reading_minutes: 5
citation: "Sebastian Raschka, Sebastian Raschka. \"Language Models for Text Classification: From Bag-of-Words to Jev.\" 29 Sept 2026. https://magazine.sebastianraschka.com/p/classifier-history-and-jev (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# Language Models for Text Classification: From Bag-of-Words to Jev

*A visual guide to bag-of-words, RNNs, CNNs, transformers, Jev-like APIs, and calibration*

> Sebastian Raschka walks from classic bag-of-words classifiers through RNNs, CNNs, and transformers to TypeSafe AI's Jev—explaining APIs, IMDb benchmarks, calibration, and why decision models matter for agent harnesses.

# Language Models for Text Classification: From Bag-of-Words to Jev

### A Visual Guide to Bag-of-Words, RNNs, CNNs, Transformers, Jev-like APIs, and Calibration

*Sebastian Raschka, PhD — Sep 29, 2026*

The recently released Jev AI model has been quite a cultural phenomenon in technical communities in the past 2 weeks.

While Jev aims to classify things, it's easy to dismiss Jev as "just a classifier," and my own view of Jev has evolved quite a bit over the past few days. In particular, my thoughts went from "classifiers used to be my bread & butter; I can easily build this myself" to "wow, this actually works better than I thought."

Sure, the latest state-of-the-art GPT and open-weight LLMs can do the same kinds of classification tasks as Jev, while also being capable of much more general decision-making. But Jev's advantage is that it can handle those classification tasks much faster and more cheaply.

At the other end of the spectrum, for a narrow, well-defined problem, Jev probably won't classify anything better, faster, or cheaper than a special-purpose classifier. But its selling point is that it is far more general than those task-specific models.

PS: I am not affiliated with Jev in any way. Also, I am not offered free access to Jev, and this is also not a product endorsement, just a technical article to offer some insights into the history of text classification to help you make sense of the recent hype.

## 1. Language modeling and classification in the pre-transformer era

### 1.1 Bag-of-words: naive Bayes, logistic regression, and XGBoost

Back in the day, when I was a grad student 15 years ago, even though recurrent neural networks already existed, text classification was usually done with a bag-of-words representation because it was straightforward and could get good results on moderately sized datasets.

In short, we can think of the bag-of-words representation as a method that makes free-form text input of different lengths compatible with classic classifiers (naive Bayes, Logistic Regression, SVMs, Random Forest, XGBoost), which expect a fixed-size input vector.

Popular real-world applications include anything from news article classification to email spam filtering. And yes, allegedly even Gmail's original spam filter used a Naive Bayes model with a bag-of-words representation.

A bag-of-words model starts by building the vocabulary, which consists of all unique words in the training set. It assigns each word in a vocabulary its own position in a vector and counts how often each word occurs in a document. One of the biggest downsides is that it loses word order — "the dog bites the man" and "the man bites the dog" produce identical vectors.

Despite the shortcomings, a bag-of-words representation + logistic regression remains my go-to baseline for every text classification problem, since it's so easy to implement. On IMDb movie reviews, this model achieves about 89.9% accuracy (on a balanced dataset).

### 1.2 Deep neural networks for text classification

More sophisticated neural network architectures avoid the bag-of-words workaround: convolutional neural networks (CNNs) and recurrent neural networks (RNNs), which can take word embeddings as input.

**Word embeddings** convert individual words into dense vectors of learned numbers (Word2Vec, GloVe, or embedding layers). Classic embeddings are context-independent at lookup time.

**RNNs** read a sequence one word at a time, combining the current word embedding with a hidden state from the previous step. LSTMs (1997) and GRUs (2014) improved trainability. A simple LSTM trained from scratch on IMDb achieved about 85.66% accuracy; ULMFiT (2018) transfer learning reached 95.4% test accuracy.

**CNNs for text** apply learned filters to windows of adjacent word embeddings. Parallel filter computations avoid RNN step-by-step dependency. On IMDb, such a CNN got about 90.07% accuracy in the author's experiments.

## 2. Transformers

In 2017, the original transformer architecture was introduced in *Attention Is All You Need*. The original was an encoder-decoder for translation, but it can be adapted for text classification.

### 2.1 Encoder-style language models

Encoder-style models like BERT were natural text classifiers. BERT has a classification token at the first position that can be fine-tuned. ModernBERT (2024) gets approximately 95% accuracy on IMDb with very little fine-tuning effort.

### 2.2 Decoder-style LLMs

GPT-style models can be prompted for classification, but for structured outputs and a specific target domain this is unnecessarily brittle and inefficient. Instead, replace the output layer with a leaner classification head. A relatively small GPT-2 124M model gets approximately 92% accuracy on IMDb.

### 2.3 Encoder-decoder style architectures

T5 (2019) can use either a classification head or text-to-text classification (decoder outputs the class label).

## 3. Jev overview

Jev is a proprietary model from TypeSafe AI that claims to be on par with GPT-5.6 Luna for decision-making while being orders of magnitude faster and cheaper. One reason the tech community is excited is that it is the "ChatGPT moment for classification": it can cheaply classify all kinds of text inputs without custom fine-tuning for each task.

### 3.2 The Jev API

Three main API types:

- **Choice API** — multi-class classification with explicit criteria
- **Noul API** — binary / multi-label "yes" probabilities
- **Score API** — ordinal classification against a rubric

### 3.3 Jev classifying IMDb

Running Jev Choice on the 25,000-review IMDb test set: **96.47% accuracy** in ~22 minutes for about $0.65. Noul: **96.20%**. A fine-tuned ModernBERT had similar accuracy but can't play Tetris or generalize to arbitrary new tasks without re-fine-tuning.

## 4–5. Retrofitting a Jev-like API; architecture and calibration

Adding a Jev-like Choice API on top of ModernBERT or GPT is straightforward: a one-node scoring head over candidate descriptions, with softmax across candidates. Matching Jev's breadth across tasks (including Tetris) is not trivial — most quick clones fall short.

TypeSafe AI describes training via **Reinforcement Learning for Calibrated Decisions (RLCD)** on 100% synthetic data. A related public method is **RLCR** (*Beyond Binary Rewards*, 2025), which adds a Brier-style penalty for inaccurate confidence alongside answer correctness. Calibration matters in production whenever you act on class-membership probabilities.

## 6. Who is Jev for?

Use cases include spam/prioritization filters and augmenting LLM-agent harnesses (prompt-injection pre-screening, effort selection, skill routing, judges, context file selection). OpenAI announced a similar Decisions API at DevDay 2026.

## Conclusion

At first glance, Jev doesn't seem to offer anything fundamentally new — "it's just a classifier" with a nice API. But it works surprisingly well across a huge range of tasks. Experts will still fine-tune specialists, but the bar to justify fine-tuning is now much higher. Making such models part of agent harnesses could make those harnesses much faster and cheaper.

Thanks for reading and supporting my independent research!

*Full original with figures: [magazine.sebastianraschka.com/p/classifier-history-and-jev](https://magazine.sebastianraschka.com/p/classifier-history-and-jev)*
