Language Models for Text Classification: From Bag-of-Words to Jev
A Visual Guide to Bag-of-Words, RNNs, CNNs, Transformers, Jev-like APIs, and Calibration
Sebastian Raschka, PhD — Sep 29, 2026
The recently released Jev AI model has been quite a cultural phenomenon in technical communities in the past 2 weeks.
While Jev aims to classify things, it's easy to dismiss Jev as "just a classifier," and my own view of Jev has evolved quite a bit over the past few days. In particular, my thoughts went from "classifiers used to be my bread & butter; I can easily build this myself" to "wow, this actually works better than I thought."
Sure, the latest state-of-the-art GPT and open-weight LLMs can do the same kinds of classification tasks as Jev, while also being capable of much more general decision-making. But Jev's advantage is that it can handle those classification tasks much faster and more cheaply.
At the other end of the spectrum, for a narrow, well-defined problem, Jev probably won't classify anything better, faster, or cheaper than a special-purpose classifier. But its selling point is that it is far more general than those task-specific models.
PS: I am not affiliated with Jev in any way. Also, I am not offered free access to Jev, and this is also not a product endorsement, just a technical article to offer some insights into the history of text classification to help you make sense of the recent hype.
1. Language modeling and classification in the pre-transformer era
1.1 Bag-of-words: naive Bayes, logistic regression, and XGBoost
Back in the day, when I was a grad student 15 years ago, even though recurrent neural networks already existed, text classification was usually done with a bag-of-words representation because it was straightforward and could get good results on moderately sized datasets.
In short, we can think of the bag-of-words representation as a method that makes free-form text input of different lengths compatible with classic classifiers (naive Bayes, Logistic Regression, SVMs, Random Forest, XGBoost), which expect a fixed-size input vector.
Popular real-world applications include anything from news article classification to email spam filtering. And yes, allegedly even Gmail's original spam filter used a Naive Bayes model with a bag-of-words representation.
A bag-of-words model starts by building the vocabulary, which consists of all unique words in the training set. It assigns each word in a vocabulary its own position in a vector and counts how often each word occurs in a document. One of the biggest downsides is that it loses word order — "the dog bites the man" and "the man bites the dog" produce identical vectors.
Despite the shortcomings, a bag-of-words representation + logistic regression remains my go-to baseline for every text classification problem, since it's so easy to implement. On IMDb movie reviews, this model achieves about 89.9% accuracy (on a balanced dataset).
1.2 Deep neural networks for text classification
More sophisticated neural network architectures avoid the bag-of-words workaround: convolutional neural networks (CNNs) and recurrent neural networks (RNNs), which can take word embeddings as input.
Word embeddings convert individual words into dense vectors of learned numbers (Word2Vec, GloVe, or embedding layers). Classic embeddings are context-independent at lookup time.
RNNs read a sequence one word at a time, combining the current word embedding with a hidden state from the previous step. LSTMs (1997) and GRUs (2014) improved trainability. A simple LSTM trained from scratch on IMDb achieved about 85.66% accuracy; ULMFiT (2018) transfer learning reached 95.4% test accuracy.
CNNs for text apply learned filters to windows of adjacent word embeddings. Parallel filter computations avoid RNN step-by-step dependency. On IMDb, such a CNN got about 90.07% accuracy in the author's experiments.
2. Transformers
In 2017, the original transformer architecture was introduced in Attention Is All You Need. The original was an encoder-decoder for translation, but it can be adapted for text classification.
2.1 Encoder-style language models
Encoder-style models like BERT were natural text classifiers. BERT has a classification token at the first position that can be fine-tuned. ModernBERT (2024) gets approximately 95% accuracy on IMDb with very little fine-tuning effort.
2.2 Decoder-style LLMs
GPT-style models can be prompted for classification, but for structured outputs and a specific target domain this is unnecessarily brittle and inefficient. Instead, replace the output layer with a leaner classification head. A relatively small GPT-2 124M model gets approximately 92% accuracy on IMDb.
2.3 Encoder-decoder style architectures
T5 (2019) can use either a classification head or text-to-text classification (decoder outputs the class label).
3. Jev overview
Jev is a proprietary model from TypeSafe AI that claims to be on par with GPT-5.6 Luna for decision-making while being orders of magnitude faster and cheaper. One reason the tech community is excited is that it is the "ChatGPT moment for classification": it can cheaply classify all kinds of text inputs without custom fine-tuning for each task.
3.2 The Jev API
Three main API types:
- Choice API — multi-class classification with explicit criteria
- Noul API — binary / multi-label "yes" probabilities
- Score API — ordinal classification against a rubric
3.3 Jev classifying IMDb
Running Jev Choice on the 25,000-review IMDb test set: 96.47% accuracy in ~22 minutes for about $0.65. Noul: 96.20%. A fine-tuned ModernBERT had similar accuracy but can't play Tetris or generalize to arbitrary new tasks without re-fine-tuning.
4–5. Retrofitting a Jev-like API; architecture and calibration
Adding a Jev-like Choice API on top of ModernBERT or GPT is straightforward: a one-node scoring head over candidate descriptions, with softmax across candidates. Matching Jev's breadth across tasks (including Tetris) is not trivial — most quick clones fall short.
TypeSafe AI describes training via Reinforcement Learning for Calibrated Decisions (RLCD) on 100% synthetic data. A related public method is RLCR (Beyond Binary Rewards, 2025), which adds a Brier-style penalty for inaccurate confidence alongside answer correctness. Calibration matters in production whenever you act on class-membership probabilities.
6. Who is Jev for?
Use cases include spam/prioritization filters and augmenting LLM-agent harnesses (prompt-injection pre-screening, effort selection, skill routing, judges, context file selection). OpenAI announced a similar Decisions API at DevDay 2026.
Conclusion
At first glance, Jev doesn't seem to offer anything fundamentally new — "it's just a classifier" with a nice API. But it works surprisingly well across a huge range of tasks. Experts will still fine-tune specialists, but the bar to justify fine-tuning is now much higher. Making such models part of agent harnesses could make those harnesses much faster and cheaper.
Thanks for reading and supporting my independent research!
Full original with figures: magazine.sebastianraschka.com/p/classifier-history-and-jev