| 1. | | GPT-6 Astra, looped transformers, and hidden reasoning (sebastianraschka.com) |
| 454 points by ModelForge 21 hours ago | past | 146 comments |
|
| 2. | | Claude Watermarks Text: Token sampling, watermark detection, and removal (sebastianraschka.com) |
| 3 points by ModelForge 16 days ago | past |
|
| 3. | | How Claude's Text Watermarking Works [video] (youtube.com) |
| 2 points by ModelForge 21 days ago | past |
|
| 4. | | How Claude's Text Watermarking Works (sebastianraschka.com) |
| 5 points by ModelForge 23 days ago | past |
|
| 5. | | Kimi K3 Architecture Overview and Notes (sebastianraschka.com) |
| 507 points by ModelForge 43 days ago | past | 111 comments |
|
| 6. | | Inkling: A New Open-Weight 975B Moe with a Few Surprises (sebastianraschka.com) |
| 3 points by ModelForge 55 days ago | past |
|
| 7. | | Claude Code's Real Secret Sauce Isn't the Model (sebastianraschka.com) |
| 6 points by ModelForge 5 months ago | past |
|
| 8. | | The State of LLMs 2025: Progress, Problems, and Predictions (sebastianraschka.com) |
| 3 points by ModelForge 8 months ago | past |
|
| 9. | | A Researcher's Field Guide to Non-Standard LLM Architectures (sebastianraschka.com) |
| 2 points by ModelForge 10 months ago | past |
|
| 10. | | Explanation of Gated DeltaNet (Qwen3-Next and Kimi Linear) (github.com/rasbt) |
| 3 points by ModelForge 10 months ago | past |
|
| 11. | | The Core Components of Modern LLMs and the Models Beyond Transformers [video] (youtube.com) |
| 3 points by ModelForge 10 months ago | past |
|
| 12. | | Popular Attention Alternatives: GQA, MLA, SWA (sebastianraschka.com) |
| 4 points by ModelForge 10 months ago | past |
|
| 13. | | Multi-Head Latent Attention (sebastianraschka.com) |
| 4 points by ModelForge 11 months ago | past |
|
| 14. | | Thinking Machines Lab Co-Founder Departs for Meta (wsj.com) |
| 7 points by ModelForge 11 months ago | past |
|
| 15. | | OpenAI's internal Slack messages could cost it billions in copyright suit (sherwood.news) |
| 8 points by ModelForge 11 months ago | past | 1 comment |
|
| 16. | | LLM Evaluation from Scratch: Multiple Choice, Verifiers, Leaderboards, LLM Judge (sebastianraschka.com) |
| 4 points by ModelForge 11 months ago | past |
|
| 17. | | Gemma 3 270M re-implemented in pure PyTorch for local tinkering (github.com/rasbt) |
| 417 points by ModelForge on Aug 20, 2025 | past | 57 comments |
|
| 18. | | GPT-OSS vs. Qwen3 and a detailed look how things evolved since GPT-2 (sebastianraschka.com) |
| 490 points by ModelForge on Aug 10, 2025 | past | 97 comments |
|
| 19. | | LLM Research Papers: The 2024 List (sebastianraschka.com) |
| 5 points by ModelForge on Dec 18, 2024 | past |
|
| 20. | | Scaling Test-Time Compute with Open LLM Models (huggingface.co) |
| 3 points by ModelForge on Dec 18, 2024 | past |
|