Language Models for Text Classification: From Bag-of-Words to Jev
A Visual Guide to RNNs, CNNs, Transformers, and Calibration, with Hands-On Experiments on Accuracy and Efficiency


Newsletter by Sebastian Raschka
Ahead of AI focuses on machine learning and AI research and is read by more than 200,000 researchers and practitioners who want to stay ahead in a rapidly evolving field.
5 posts in the last 90 days · newest 29 Sept
A Visual Guide to RNNs, CNNs, Transformers, and Calibration, with Hands-On Experiments on Accuracy and Efficiency

A Look at Recurrent Depth, Hidden Chains of Thought, and Recent Research on Looping Transformer Blocks

A 48-minute video walkthrough of token sampling, watermark detection, and removal

An End-to-End Project With Dataset Construction, Model Training, Local Deployment, and RLVR

How LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes

Using Open-Weight Models in Local Coding Harnesses as an Alternative to Claude Code and Codex Subscriptions

A curated roundup of notable LLM research papers that came out this year

From Gemma 4 to DeepSeek V4, How New Open-Weight LLMs Are Reducing Long-Context Costs

A learning-oriented workflow for understanding new open-weight model releases

How coding agents use tools, memory, and repo context to make LLMs work better in practice

From MHA and GQA to MLA, sparse attention, and hybrid architectures

A Round Up And Comparison of 10 Open-Weight LLM Releases in Spring 2026

And an Overview of Recent Inference-Scaling Papers

A 2025 review of large language models, from DeepSeek R1 and RLVR to inference-time scaling, benchmarks, architectures, and predictions for 2026.

In June, I shared a bonus article with my curated and bookmarked research paper lists to the paid subscribers who make this Substack possible.

Understanding How DeepSeek's Flagship Open-Weight Models Evolved

Linear Attention Hybrids, Text Diffusion, Code World Models, and Small Recursive Transformers
