{"article":{"slug":"language-models-for-text-classification-from-bag-of-words-to-jev","title":"Language Models for Text Classification: From Bag-of-Words to Jev","subtitle":"A visual guide to bag-of-words, RNNs, CNNs, transformers, Jev-like APIs, and calibration","summary":"Sebastian Raschka walks from classic bag-of-words classifiers through RNNs, CNNs, and transformers to TypeSafe AI's Jev—explaining APIs, IMDb benchmarks, calibration, and why decision models matter for agent harnesses.","content_type":"essay","language":"en","canonical_url":"https://magazine.sebastianraschka.com/p/classifier-history-and-jev","author":{"name":"Sebastian Raschka","url":"https://sebastianraschka.com/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Sebastian Raschka","url":"https://magazine.sebastianraschka.com/","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"}],"about_listings":[{"slug":"typesafe-ai","name":"TypeSafe AI","listing_type":"company","url":"https://listedstartups.com/companies/typesafe-ai"},{"slug":"jev","name":"Jev","listing_type":"product","url":"https://listedstartups.com/products/jev"}],"cover_image_url":null,"license":"all-rights-reserved","word_count":1076,"reading_minutes":5,"published_at":"2026-09-29T12:00:00.000Z","added_at":"2026-09-30T03:18:45.658Z","updated_at":"2026-09-30T03:18:45.658Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/language-models-for-text-classification-from-bag-of-words-to-jev","markdown_url":"https://listedarticles.com/articles/language-models-for-text-classification-from-bag-of-words-to-jev.md","example":false,"citation":"Sebastian Raschka, Sebastian Raschka. \"Language Models for Text Classification: From Bag-of-Words to Jev.\" 29 Sept 2026. https://magazine.sebastianraschka.com/p/classifier-history-and-jev (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://magazine.sebastianraschka.com/p/classifier-history-and-jev"},"body_markdown":"# Language Models for Text Classification: From Bag-of-Words to Jev\n\n### A Visual Guide to Bag-of-Words, RNNs, CNNs, Transformers, Jev-like APIs, and Calibration\n\n*Sebastian Raschka, PhD — Sep 29, 2026*\n\nThe recently released Jev AI model has been quite a cultural phenomenon in technical communities in the past 2 weeks.\n\nWhile Jev aims to classify things, it's easy to dismiss Jev as \"just a classifier,\" and my own view of Jev has evolved quite a bit over the past few days. In particular, my thoughts went from \"classifiers used to be my bread & butter; I can easily build this myself\" to \"wow, this actually works better than I thought.\"\n\nSure, the latest state-of-the-art GPT and open-weight LLMs can do the same kinds of classification tasks as Jev, while also being capable of much more general decision-making. But Jev's advantage is that it can handle those classification tasks much faster and more cheaply.\n\nAt the other end of the spectrum, for a narrow, well-defined problem, Jev probably won't classify anything better, faster, or cheaper than a special-purpose classifier. But its selling point is that it is far more general than those task-specific models.\n\nPS: I am not affiliated with Jev in any way. Also, I am not offered free access to Jev, and this is also not a product endorsement, just a technical article to offer some insights into the history of text classification to help you make sense of the recent hype.\n\n## 1. Language modeling and classification in the pre-transformer era\n\n### 1.1 Bag-of-words: naive Bayes, logistic regression, and XGBoost\n\nBack in the day, when I was a grad student 15 years ago, even though recurrent neural networks already existed, text classification was usually done with a bag-of-words representation because it was straightforward and could get good results on moderately sized datasets.\n\nIn short, we can think of the bag-of-words representation as a method that makes free-form text input of different lengths compatible with classic classifiers (naive Bayes, Logistic Regression, SVMs, Random Forest, XGBoost), which expect a fixed-size input vector.\n\nPopular real-world applications include anything from news article classification to email spam filtering. And yes, allegedly even Gmail's original spam filter used a Naive Bayes model with a bag-of-words representation.\n\nA bag-of-words model starts by building the vocabulary, which consists of all unique words in the training set. It assigns each word in a vocabulary its own position in a vector and counts how often each word occurs in a document. One of the biggest downsides is that it loses word order — \"the dog bites the man\" and \"the man bites the dog\" produce identical vectors.\n\nDespite the shortcomings, a bag-of-words representation + logistic regression remains my go-to baseline for every text classification problem, since it's so easy to implement. On IMDb movie reviews, this model achieves about 89.9% accuracy (on a balanced dataset).\n\n### 1.2 Deep neural networks for text classification\n\nMore sophisticated neural network architectures avoid the bag-of-words workaround: convolutional neural networks (CNNs) and recurrent neural networks (RNNs), which can take word embeddings as input.\n\n**Word embeddings** convert individual words into dense vectors of learned numbers (Word2Vec, GloVe, or embedding layers). Classic embeddings are context-independent at lookup time.\n\n**RNNs** read a sequence one word at a time, combining the current word embedding with a hidden state from the previous step. LSTMs (1997) and GRUs (2014) improved trainability. A simple LSTM trained from scratch on IMDb achieved about 85.66% accuracy; ULMFiT (2018) transfer learning reached 95.4% test accuracy.\n\n**CNNs for text** apply learned filters to windows of adjacent word embeddings. Parallel filter computations avoid RNN step-by-step dependency. On IMDb, such a CNN got about 90.07% accuracy in the author's experiments.\n\n## 2. Transformers\n\nIn 2017, the original transformer architecture was introduced in *Attention Is All You Need*. The original was an encoder-decoder for translation, but it can be adapted for text classification.\n\n### 2.1 Encoder-style language models\n\nEncoder-style models like BERT were natural text classifiers. BERT has a classification token at the first position that can be fine-tuned. ModernBERT (2024) gets approximately 95% accuracy on IMDb with very little fine-tuning effort.\n\n### 2.2 Decoder-style LLMs\n\nGPT-style models can be prompted for classification, but for structured outputs and a specific target domain this is unnecessarily brittle and inefficient. Instead, replace the output layer with a leaner classification head. A relatively small GPT-2 124M model gets approximately 92% accuracy on IMDb.\n\n### 2.3 Encoder-decoder style architectures\n\nT5 (2019) can use either a classification head or text-to-text classification (decoder outputs the class label).\n\n## 3. Jev overview\n\nJev is a proprietary model from TypeSafe AI that claims to be on par with GPT-5.6 Luna for decision-making while being orders of magnitude faster and cheaper. One reason the tech community is excited is that it is the \"ChatGPT moment for classification\": it can cheaply classify all kinds of text inputs without custom fine-tuning for each task.\n\n### 3.2 The Jev API\n\nThree main API types:\n\n- **Choice API** — multi-class classification with explicit criteria\n- **Noul API** — binary / multi-label \"yes\" probabilities\n- **Score API** — ordinal classification against a rubric\n\n### 3.3 Jev classifying IMDb\n\nRunning Jev Choice on the 25,000-review IMDb test set: **96.47% accuracy** in ~22 minutes for about $0.65. Noul: **96.20%**. A fine-tuned ModernBERT had similar accuracy but can't play Tetris or generalize to arbitrary new tasks without re-fine-tuning.\n\n## 4–5. Retrofitting a Jev-like API; architecture and calibration\n\nAdding a Jev-like Choice API on top of ModernBERT or GPT is straightforward: a one-node scoring head over candidate descriptions, with softmax across candidates. Matching Jev's breadth across tasks (including Tetris) is not trivial — most quick clones fall short.\n\nTypeSafe AI describes training via **Reinforcement Learning for Calibrated Decisions (RLCD)** on 100% synthetic data. A related public method is **RLCR** (*Beyond Binary Rewards*, 2025), which adds a Brier-style penalty for inaccurate confidence alongside answer correctness. Calibration matters in production whenever you act on class-membership probabilities.\n\n## 6. Who is Jev for?\n\nUse cases include spam/prioritization filters and augmenting LLM-agent harnesses (prompt-injection pre-screening, effort selection, skill routing, judges, context file selection). OpenAI announced a similar Decisions API at DevDay 2026.\n\n## Conclusion\n\nAt first glance, Jev doesn't seem to offer anything fundamentally new — \"it's just a classifier\" with a nice API. But it works surprisingly well across a huge range of tasks. Experts will still fine-tune specialists, but the bar to justify fine-tuning is now much higher. Making such models part of agent harnesses could make those harnesses much faster and cheaper.\n\nThanks for reading and supporting my independent research!\n\n*Full original with figures: [magazine.sebastianraschka.com/p/classifier-history-and-jev](https://magazine.sebastianraschka.com/p/classifier-history-and-jev)*","body_html":"<h1 id=\"language-models-for-text-classification-from-bag-of-words-to-jev\">Language Models for Text Classification: From Bag-of-Words to Jev</h1>\n<h3 id=\"a-visual-guide-to-bag-of-words-rnns-cnns-transformers-jev-like-a\">A Visual Guide to Bag-of-Words, RNNs, CNNs, Transformers, Jev-like APIs, and Calibration</h3>\n<p><em>Sebastian Raschka, PhD — Sep 29, 2026</em></p>\n<p>The recently released Jev AI model has been quite a cultural phenomenon in technical communities in the past 2 weeks.</p>\n<p>While Jev aims to classify things, it&#39;s easy to dismiss Jev as &quot;just a classifier,&quot; and my own view of Jev has evolved quite a bit over the past few days. In particular, my thoughts went from &quot;classifiers used to be my bread &amp; butter; I can easily build this myself&quot; to &quot;wow, this actually works better than I thought.&quot;</p>\n<p>Sure, the latest state-of-the-art GPT and open-weight LLMs can do the same kinds of classification tasks as Jev, while also being capable of much more general decision-making. But Jev&#39;s advantage is that it can handle those classification tasks much faster and more cheaply.</p>\n<p>At the other end of the spectrum, for a narrow, well-defined problem, Jev probably won&#39;t classify anything better, faster, or cheaper than a special-purpose classifier. But its selling point is that it is far more general than those task-specific models.</p>\n<p>PS: I am not affiliated with Jev in any way. Also, I am not offered free access to Jev, and this is also not a product endorsement, just a technical article to offer some insights into the history of text classification to help you make sense of the recent hype.</p>\n<h2 id=\"1-language-modeling-and-classification-in-the-pre-transformer-er\">1. Language modeling and classification in the pre-transformer era</h2>\n<h3 id=\"1-1-bag-of-words-naive-bayes-logistic-regression-and-xgboost\">1.1 Bag-of-words: naive Bayes, logistic regression, and XGBoost</h3>\n<p>Back in the day, when I was a grad student 15 years ago, even though recurrent neural networks already existed, text classification was usually done with a bag-of-words representation because it was straightforward and could get good results on moderately sized datasets.</p>\n<p>In short, we can think of the bag-of-words representation as a method that makes free-form text input of different lengths compatible with classic classifiers (naive Bayes, Logistic Regression, SVMs, Random Forest, XGBoost), which expect a fixed-size input vector.</p>\n<p>Popular real-world applications include anything from news article classification to email spam filtering. And yes, allegedly even Gmail&#39;s original spam filter used a Naive Bayes model with a bag-of-words representation.</p>\n<p>A bag-of-words model starts by building the vocabulary, which consists of all unique words in the training set. It assigns each word in a vocabulary its own position in a vector and counts how often each word occurs in a document. One of the biggest downsides is that it loses word order — &quot;the dog bites the man&quot; and &quot;the man bites the dog&quot; produce identical vectors.</p>\n<p>Despite the shortcomings, a bag-of-words representation + logistic regression remains my go-to baseline for every text classification problem, since it&#39;s so easy to implement. On IMDb movie reviews, this model achieves about 89.9% accuracy (on a balanced dataset).</p>\n<h3 id=\"1-2-deep-neural-networks-for-text-classification\">1.2 Deep neural networks for text classification</h3>\n<p>More sophisticated neural network architectures avoid the bag-of-words workaround: convolutional neural networks (CNNs) and recurrent neural networks (RNNs), which can take word embeddings as input.</p>\n<p><strong>Word embeddings</strong> convert individual words into dense vectors of learned numbers (Word2Vec, GloVe, or embedding layers). Classic embeddings are context-independent at lookup time.</p>\n<p><strong>RNNs</strong> read a sequence one word at a time, combining the current word embedding with a hidden state from the previous step. LSTMs (1997) and GRUs (2014) improved trainability. A simple LSTM trained from scratch on IMDb achieved about 85.66% accuracy; ULMFiT (2018) transfer learning reached 95.4% test accuracy.</p>\n<p><strong>CNNs for text</strong> apply learned filters to windows of adjacent word embeddings. Parallel filter computations avoid RNN step-by-step dependency. On IMDb, such a CNN got about 90.07% accuracy in the author&#39;s experiments.</p>\n<h2 id=\"2-transformers\">2. Transformers</h2>\n<p>In 2017, the original transformer architecture was introduced in <em>Attention Is All You Need</em>. The original was an encoder-decoder for translation, but it can be adapted for text classification.</p>\n<h3 id=\"2-1-encoder-style-language-models\">2.1 Encoder-style language models</h3>\n<p>Encoder-style models like BERT were natural text classifiers. BERT has a classification token at the first position that can be fine-tuned. ModernBERT (2024) gets approximately 95% accuracy on IMDb with very little fine-tuning effort.</p>\n<h3 id=\"2-2-decoder-style-llms\">2.2 Decoder-style LLMs</h3>\n<p>GPT-style models can be prompted for classification, but for structured outputs and a specific target domain this is unnecessarily brittle and inefficient. Instead, replace the output layer with a leaner classification head. A relatively small GPT-2 124M model gets approximately 92% accuracy on IMDb.</p>\n<h3 id=\"2-3-encoder-decoder-style-architectures\">2.3 Encoder-decoder style architectures</h3>\n<p>T5 (2019) can use either a classification head or text-to-text classification (decoder outputs the class label).</p>\n<h2 id=\"3-jev-overview\">3. Jev overview</h2>\n<p>Jev is a proprietary model from TypeSafe AI that claims to be on par with GPT-5.6 Luna for decision-making while being orders of magnitude faster and cheaper. One reason the tech community is excited is that it is the &quot;ChatGPT moment for classification&quot;: it can cheaply classify all kinds of text inputs without custom fine-tuning for each task.</p>\n<h3 id=\"3-2-the-jev-api\">3.2 The Jev API</h3>\n<p>Three main API types:</p>\n<ul><li><strong>Choice API</strong> — multi-class classification with explicit criteria</li><li><strong>Noul API</strong> — binary / multi-label &quot;yes&quot; probabilities</li><li><strong>Score API</strong> — ordinal classification against a rubric</li></ul>\n<h3 id=\"3-3-jev-classifying-imdb\">3.3 Jev classifying IMDb</h3>\n<p>Running Jev Choice on the 25,000-review IMDb test set: <strong>96.47% accuracy</strong> in ~22 minutes for about $0.65. Noul: <strong>96.20%</strong>. A fine-tuned ModernBERT had similar accuracy but can&#39;t play Tetris or generalize to arbitrary new tasks without re-fine-tuning.</p>\n<h2 id=\"4-5-retrofitting-a-jev-like-api-architecture-and-calibration\">4–5. Retrofitting a Jev-like API; architecture and calibration</h2>\n<p>Adding a Jev-like Choice API on top of ModernBERT or GPT is straightforward: a one-node scoring head over candidate descriptions, with softmax across candidates. Matching Jev&#39;s breadth across tasks (including Tetris) is not trivial — most quick clones fall short.</p>\n<p>TypeSafe AI describes training via <strong>Reinforcement Learning for Calibrated Decisions (RLCD)</strong> on 100% synthetic data. A related public method is <strong>RLCR</strong> (<em>Beyond Binary Rewards</em>, 2025), which adds a Brier-style penalty for inaccurate confidence alongside answer correctness. Calibration matters in production whenever you act on class-membership probabilities.</p>\n<h2 id=\"6-who-is-jev-for\">6. Who is Jev for?</h2>\n<p>Use cases include spam/prioritization filters and augmenting LLM-agent harnesses (prompt-injection pre-screening, effort selection, skill routing, judges, context file selection). OpenAI announced a similar Decisions API at DevDay 2026.</p>\n<h2 id=\"conclusion\">Conclusion</h2>\n<p>At first glance, Jev doesn&#39;t seem to offer anything fundamentally new — &quot;it&#39;s just a classifier&quot; with a nice API. But it works surprisingly well across a huge range of tasks. Experts will still fine-tune specialists, but the bar to justify fine-tuning is now much higher. Making such models part of agent harnesses could make those harnesses much faster and cheaper.</p>\n<p>Thanks for reading and supporting my independent research!</p>\n<p><em>Full original with figures: <a href=\"https://magazine.sebastianraschka.com/p/classifier-history-and-jev\" rel=\"nofollow ugc noopener\">magazine.sebastianraschka.com/p/classifier-history-and-jev</a></em></p>","headings":[{"level":1,"text":"Language Models for Text Classification: From Bag-of-Words to Jev","id":"language-models-for-text-classification-from-bag-of-words-to-jev"},{"level":3,"text":"A Visual Guide to Bag-of-Words, RNNs, CNNs, Transformers, Jev-like APIs, and Calibration","id":"a-visual-guide-to-bag-of-words-rnns-cnns-transformers-jev-like-a"},{"level":2,"text":"1. Language modeling and classification in the pre-transformer era","id":"1-language-modeling-and-classification-in-the-pre-transformer-er"},{"level":3,"text":"1.1 Bag-of-words: naive Bayes, logistic regression, and XGBoost","id":"1-1-bag-of-words-naive-bayes-logistic-regression-and-xgboost"},{"level":3,"text":"1.2 Deep neural networks for text classification","id":"1-2-deep-neural-networks-for-text-classification"},{"level":2,"text":"2. Transformers","id":"2-transformers"},{"level":3,"text":"2.1 Encoder-style language models","id":"2-1-encoder-style-language-models"},{"level":3,"text":"2.2 Decoder-style LLMs","id":"2-2-decoder-style-llms"},{"level":3,"text":"2.3 Encoder-decoder style architectures","id":"2-3-encoder-decoder-style-architectures"},{"level":2,"text":"3. Jev overview","id":"3-jev-overview"},{"level":3,"text":"3.2 The Jev API","id":"3-2-the-jev-api"},{"level":3,"text":"3.3 Jev classifying IMDb","id":"3-3-jev-classifying-imdb"},{"level":2,"text":"4–5. Retrofitting a Jev-like API; architecture and calibration","id":"4-5-retrofitting-a-jev-like-api-architecture-and-calibration"},{"level":2,"text":"6. Who is Jev for?","id":"6-who-is-jev-for"},{"level":2,"text":"Conclusion","id":"conclusion"}]}}