# Aadit Suryawanshi > I train and fine-tune language models, build memory and retrieval systems, and ship full-stack AI products on scalable backend infrastructure. Site: https://aadit032.github.io/portfolio/ _Auto-generated agent context from the portfolio. Prefer this file over scraping HTML._ ## How to use this file This is the full context dump for Aadit Suryawanshi’s portfolio. It includes profile, skills focus, projects (full text), and blog articles (full text). Use it when answering questions about who I am, what I build, or the work on this site. ## About Aadit Suryawanshi builds across the AI stack end-to-end. I can **train models from scratch** (decoder-only Transformers in PyTorch, custom tokenizers, data mixtures, training loops) and **fine-tune open-source LLMs** (continued pre-training, LoRA/SFT, domain adaptation, evaluation). I also work on problems **beyond the model weights**: - **Memory & retrieval** — multimodal ingestion, hybrid search (dense + sparse), reranking, agent memory architectures that persist knowledge across conversations and documents - **Scalable infrastructure** — async pipelines with Redis Streams, worker systems, queues/DLQs, streaming responses, observability - **Full-stack product work** — frontend apps, APIs, auth, realtime collaboration, and wiring ML into usable systems I learn by shipping. Projects usually start as a technical question and turn into systems that combine ML, distributed backends, and product engineering. ## What I work on - Training small language models from first principles (architecture, data, tokenizer, training, eval) - Fine-tuning and domain adaptation of open-source LLMs (CPT, SFT, LoRA, medical/domain data) - Memory architectures and retrieval pipelines (hybrid search, rerank, multimodal ingest, agent memory) - Backend infrastructure for AI workloads (Redis Streams, workers, job queues, streaming chat) - Full-stack delivery (React/Next.js, Node/Bun/Express, TypeScript, realtime WebSockets, canvas UIs) ## Current focus - AI agent memory: how systems remember, retrieve, and reason over knowledge across chats and docs - Building a multimodal memory / knowledge OS (ingestion → embed → hybrid search → chat with citations) - Scalable async infra with Redis Streams and worker fleets for document processing - Fine-tuning open-source LLMs for medical QA and domain reasoning - Training GPT-style models from scratch and understanding scaling trade-offs ## Contact - **GitHub**: https://github.com/Aadit032 - **LinkedIn**: https://www.linkedin.com/in/aadit-suryawanshi-542222316 - **Hugging Face**: https://huggingface.co/Aadit-032 - **Twitter**: https://x.com/Aadit_032 - **Email**: mailto:codexbuild.dev@gmail.com ## Site map - Home: https://aadit032.github.io/portfolio/ - Blog: https://aadit032.github.io/portfolio/blog/ - About: https://aadit032.github.io/portfolio/about/ - This context (llms.txt): https://aadit032.github.io/portfolio/llms.txt - RSS: https://aadit032.github.io/portfolio/rss.xml --- # Projects 4 projects — full write-ups below. ## LiteGPT - **URL**: https://aadit032.github.io/portfolio/projects/1-litegpt/ - **Published**: 2026-07-04 - **Summary**: A clean decoder-only Transformer language model (~25M parameters) trained from scratch on a single NVIDIA A5000 GPU. > A clean, educational decoder-only Transformer language model (~25M parameters) trained from scratch on a single NVIDIA A5000 GPU. ### Architecture ``` Input Tokens [B, T] │ ▼ ┌─────────────────────┐ │ Token Embeddings │ │ [vocab, d_model] │ └─────────────────────┘ │ │ ▼ ╔══════════════════════════════╗ ║ Transformer Block × 8 ║ ║ ║ ║ RMSNorm ║ ║ │ ║ ║ ▼ ║ ║ GQA flash Attention ║ ║ │ ║ ║ ▼ ║ ║ Residual Add ║ ║ │ ║ ║ ▼ ║ ║ RMSNorm ║ ║ │ ║ ║ ▼ ║ ║ SwiGLU MLP ║ ║ │ ║ ║ ▼ ║ ║ Residual Add ║ ╚══════════════════════════════╝ │ ▼ ┌─────────────────────┐ │ Final RMSNorm │ └─────────────────────┘ │ ▼ ┌─────────────────────┐ │ LM Head │ └─────────────────────┘ │ ▼ Logits [B,T,V] ``` | Configuration | Value | | --- | --- | | Model type | Decoder-only Transformer | | Parameters | ~24.6M | | Layers | 8 | | d_model | 448 | | Attention heads | 8 query / 4 key-value (GQA, head_dim=56) | | FFN dimension | 1152 (SwiGLU) | | Context length | 512 | | Vocabulary | 16,384 (custom BPE tokenizer) | --- The model incorporates modern LLM improvements over GPT-2: - **RoPE** positional encodings - **Grouped Query Attention (GQA)** with FlashAttention - **SwiGLU** feed-forward networks - **RMSNorm** with pre-norm residual connections - **Weight tying** between token embeddings and LM head ### Datasets Trained on ~1B tokens (custom BPE tokenizer, vocab 16,384): | Dataset | Tokens | Weight | | --- | --- | --- | | FineWeb | 300M | 60% | | TinyStories | 200M | 40% | --- Data is tokenized with a custom ByteLevel BPE tokenizer (vocab size 16,384), stored as `uint16` arrays, and split 90/10 train/validation. ### Training | Hyperparameter | Value | | --- | --- | | Effective batch size | 128 | | Sequence length | 512 | | Learning rate | 6e-4 → 6e-5 (cosine decay) | | Warmup steps | 4000 | | Max steps | 40000 | | Optimizer | AdamW (β₁=0.9, β₂=0.95) | | Precision | BF16/FP16 mixed | | Tokens seen | ~1B | --- Trained on a **NVIDIA A5000 GPU** (Runpod) using Cross Entropy Loss with cosine LR schedule and periodic checkpointing. ### Results | Metric | Value | | --- | --- | | Train Loss | 2.6853755712509155 | | Val Loss | 2.8251807928085326 | | Perplexity | 16.863993599262944 | --- ## LiteGPT Evaluation Suite **Generated:** 2026-06-26 19:48:15 ## Language modeling - **Settings:** temperature=0.8, top_k=50, max_tokens=100 ### Prompt: The capital of France is > **Output:** The capital of France is well known for its great geographical experience. The historical setting is not one of a sort of the greatest of all the times, but it's a central part of the history of the country. It's a historical town that has seen a huge shift inside this famous city of the same name. The city is known as "The Dolceus" and also known as the "The Dolceus." It's a place where a lot of people live and live, ### Prompt: Once upon a time, there lived > **Output:** Once upon a time, there lived a little girl named Lily. She was very excited because she was going on a trip to see Grandma. She helped her mom pack their luggage with clothes and toys. Then, they got in the car and drove to Grandma's house. As they drove, Lily saw something strange. It was a big, gray cloud that seemed to float away. She pointed and asked her mom what it was. Her mom said it was a mild cloud, but it was a very hot day. ### Prompt: The theory of evolution states that > **Output:** The theory of evolution states that a single model is responsible for some type of behavior, not all animals, not all human beings. It is not the same concept to decide the same. It does not matter if it is the same in a species or a family, or the exact same animal. When it is no longer possible to choose one species, it is only a matter of determining the kind of a single species. |We can try to make a simple concept based on a model, name or model (or model as ### Prompt: Machine learning is > **Output:** Machine learning is part of the same, but the reason is that each semester is different, and the number of students attending is different. This is the reason for course selection and inclusion among those at lower>level learning. So, we need to make sure we understand what the learning system is, how it is being presented and what we use to help others find their own kindergarten to finish it. So, we will get to see how each is different. We will see what our teachers and --- ## Reasoning > **Settings:** temperature=0.4, top_k=40, max_tokens=100 ### Prompt: If John has 5 apples and gives 2 to Alice, then John has > **Output:** If John has 5 apples and gives 2 to Alice, then John has 5 apples and 1 to Bob, then Alice is 3 to Bob, then Alice is 3 to Bob and Bob is 3 to Bob. A few days ago, I made a couple of apples and added them to the bottom of the apple. I think I did a pretty good job with the apple, but I didn't have a lot of apples, so I just added the apples and added the apples. I made a few more apples, but I think I would add a lot more apples ### Prompt: The next number in the sequence is: 2, 4, 8, 16 > **Output:** The next number in the sequence is: 2, 4, 8, 16, 17, 17, 17, 17, 18, 18, 18, 18, 18, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, 19, ### Prompt: Earth is to Solar System as electron is to > **Output:** Earth is to Solar System as electron is to be able to generate a high-level of energy and solar power. The solar system is designed to generate a high-level of energy, which is then used to generate a high-level of energy. Solar systems can generate electricity, which is then used to generate a high-level of energy. Solar systems are used to generate electricity, which is then used to generate electricity. Solar systems are used to generate electricity. The solar system is used to generate electricity which is then --- ## World knowledge - **Settings:** temperature=0.4, top_k=40, max_tokens=100 ### Prompt: The Internet works because > **Output:** The Internet works because it’s a platform that helps people with disabilities. What is a mobile app? A mobile app is a platform that allows people to use it to work, work, and play. It allows people to use it to work, and it allows them to use it to work. What is a mobile app? A mobile app is a platform that allows people to use it to work, work, and play. It is a platform that allows people to use it to work, work ### Prompt: The largest planet in our solar system is > **Output:** The largest planet in our solar system is the Earth's solar system. It's a great solar system for solar panels and solar panels. It's also a great solar power for solar panels and solar panels. The solar system is a great solar system for solar panels and solar panels. It's also a great solar system for solar panels and solar panels. The solar system is a great solar power system for solar panels and solar panels. It's also a great solar power system for solar panels and solar panels. The solar system is --- ## Coding - **Settings:** temperature=0.2, top_k=20, max_tokens=150 ### Prompt: for i in range(10) > **Output:** for i in range(10): - i in range(10): - i in range(10): - i in range(10): - i in range(10): - i in range(10): - i in range(10): - i in range(10): - i in range(10): - i in range(10): - i in range(10): - i in range(10): - i in range(10): - i in range(10): - i in range(10): - i in range(10): - i in range(10): - i in range(10): - i in range(10): - i in range( ### Prompt: #include int main() { > **Output:** #include int main() {\iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS/iOS ### Prompt: function quickSort(arr) { > **Output:** function quickSort(arr) {iphone(s) iphone(s) iphone(s) iphone(s) iphone(s) iphone(s) iphone(s) iphone(s) iphone(s) iphone(s) iphone(s) iphone(s) iphone(s) iphone(s) iphone(s) iphone(s) iphone(s) iphone(s) iphone(s) iphone(s) iphone(s) iphone --- ## HTML - **Settings:** temperature=0.2, top_k=20, max_tokens=150 ``` ### Prompt: **Output:** ``` --- ## Markdown - **Settings:** temperature=0.8, top_k=50, max_tokens=100 ### Prompt: # Neural Networks > **Output:** # Neural Networks. There are several networks that provide the same network as the other networks. At the same time, a network can provide a network with the same network as a wireless network (SMT). What's the difference between VoIP networks? The VoIP network has an IP address that connects both networks (i.e. IP address and IP address) and the Internet address. The network is the network's IP address that connects to a certain network network. This connects --- ## Story completion - **Settings:** temperature=0.9, top_k=100, max_tokens=200 ### Prompt: The dragon opened its eyes and > **Output:** The dragon opened its eyes and looked around. It saw something shiny and colorful. It jumped and grabbed the object with its sharp claws. The dragon flew away and the dragon breathed a sigh of joy. It flew back up into the sky and found its owner, who was so proud of it. He said the dragon had been an amazing hunter. ### Prompt: The detective walked into the abandoned warehouse and > **Output:** The detective walked into the abandoned warehouse and turned right, halfway next to the counter—the warehouse floor was all finished. I wasn’t sure what the guy was doing in front of the counter, but I was having a very slight, deliberate grimacing moment. This might be the first time I saw an account from the FBI for a meeting with a customer, or someone from the FBI’s office. This was my first time seeing the FBI’s staffer work at an incident—one that is usually a business or a place of business. My goal is to solve the problem, and there’s not as much fun as I think. A lot of people that I’m not interested in hearing from have been in the FBI’s management room, and it’s been very, very good. When I needed to tell them where it was going, I was really drawn to a number of good habits, but still had to be careful. There’s been some --- ## Dialogue - **Settings:** temperature=0.9, top_k=100, max_tokens=200 ### Prompt: Alice: Hello! Bob > **Output:** Alice: Hello! Bob: I'm sorry. Wow. You're a little bit of an understatement to me. Maybe. I know how to keep you up by yourself, even though I need to share my stories. I know, I found myself having trouble finding relationships, since I'm usually with my best friend or coworkers. I know I've been in relationship situations since I last met my friend from college and I understand I can help. I know I have found the best relationship possible for a couple of weeks now, but it's actually all about finding love. The two of you and your best friend, Mildred P. Stewart (Pennsylvania Girl). That's where the two are: a close relationship, a relationship, and an intercourse. I love it when I'm with her, and her love. These two have just been on a good track. But lately, I actually have found several relationships. "To be truly love means to be happy, --- ## TinyStories style - **Settings:** temperature=0.9, top_k=100, max_tokens=200 ### Prompt: Tom had a little red ball > **Output:** Tom had a little red ball. Lily had a red ball and Tom had a blue ball. They both loved to play with their toys. But they both wanted the red ball. They both pulled it from each other. "Give me the red ball!" Lily shouted. "It's mine!" "No, it's mine!" Tom yelled. "It's mine!" They pulled so hard that they separate the red ball from each other. They both fell down and started to cry. Mom heard the noise and came to see what was wrong. She saw the red ball on the floor and the tears on Lily's face. She was not angry, but she was sad. "Why are you fighting?" Mom asked. "You both have the red ball, but you both have the same toy." "That is boring!" Lily said. "You both need to share and take turns." "Okay!" Mom said. "But you can both have the red ball ### Prompt: Lily loved going to the park because > **Output:** Lily loved going to the park because it was always free. One day, she saw a boy crying and her knee started to hurt. 'Hello there, little girl,' the boy said, pointing to his knee. 'I'm sad because my knee hurts,' he said. The man smiled and said, 'Don't worry, I have a band-aid in my bag and I'll take it out to make you feel better.' The boy went back to his house and his mother put the band-aid on his knee. The man smiled and said, 'I'm happy to see this band-aid and I'm glad so you can help me!' The boy thanked his mother and helped the man get back in his shoes. The man said goodbye and the boy smiled as he skipped away. --- ## Long-form continuation - **Settings:** temperature=0.8, top_k=50, max_tokens=100 ### Prompt: The Great Wall of China is a series of fortifications that were built across the historical northern borders of ancient Chinese states and Imperial China as protection against various nomadic groups from the Eurasian Steppe. Several walls were built from the 7th century BC, with selective stretches later joined together by Qin Shi Huang, the first emperor of China. Little of the Qin wall remains. Later on, many successive dynasties built and maintained multiple stretches of border walls. The best-known sections of the wall were built by the Ming dynasty > **Output:** The Great Wall of China is a series of fortifications that were built across the historical northern borders of ancient Chinese states and Imperial China as protection against various nomadic groups from the Eurasian Steppe. Several walls were built from the 7th century BC, with selective stretches later joined together by Qin Shi Huang, the first emperor of China. Little of the Qin wall remains. Later on, many successive dynasties built and maintained multiple stretches of border walls. The best-known sections of the wall were built by the Ming dynasty. The Chinese Army was established in 1899, and became the first American Army occupied by the army. The city has a large population of 100,000 people; the entire city of China is the largest in the world. The city grew by 31 percent, with a total of 322,000 square feet. The most recent World War II military shrank the city's population, while the largest, and the largest, is the most populous city in China. In 1901, ## Recall-OS - **URL**: https://aadit032.github.io/portfolio/projects/recallos/ - **Published**: 2026-07-04 - **Summary**: RecallOS is an AI-native enterprise knowledge operating system that allows organizations to ingest, organize, search and reason over every piece of company knowledge. ### ✨ What is RecallOS? ![recallos](../../assets/recallos.png) A **multimodal memory architecture** for persistent retrieval over heterogeneous enterprise knowledge. Upload PDFs, images, audio, and video — then chat with an AI that cites its sources. #### 🚀 Core Features | Feature | Description | |:--------|:------------| | 📄 **Multi-modality upload** | PDF, images, audio, video via MinIO presigned URLs | | 🏭 **Modality-aware ingestion** | Per-modality parser workers dispatched by MIME type | | 🧩 **Decoupled embedding** | Modality-agnostic dense + sparse embed, re-embeddable without reparsing | | 🔍 **Hybrid search** | Dense BGE + sparse SPLADE in Qdrant, fused with RRF | | 🎯 **Cross-encoder rerank** | Top chunks reranked before LLM context injection | | 💬 **Streaming chat** | SSE streaming with source chunk citations + optional modality filter | | 🌐 **Web research agent** | `/web` prefix triggers LangGraph loop (Exa → reason → refine → answer) | | 📂 **Projects** | Organize chats with custom system prompts | | 📌 **Chat history** | Pin, delete, version (edit/resend), and rolling conversation summaries | | 📊 **Langfuse tracing** | Full observability for chat RAG and ingestion pipelines | | 🔄 **Dead Letter Queue** | Failed document processing with retry and reprocessing | | 🏗️ **Modular chat UI** | 14-file component architecture with custom hooks and focused modules | --- ### 🛠️ Tech Stack
LayerTechnology
📦 MonorepoBun workspaces + Turborepo
🖥️ FrontendNext.js 16 (App Router), React 19, Tailwind CSS v4
⚙️ BackendExpress 5 (JWT middleware on all routes except auth)
⚡ RuntimeBun
🔐 AuthJWT + bcrypt
📨 QueueRedis Streams (consumer groups, XAUTOCLAIM)
🗄️ Object storageMinIO (S3 API)
🐘 MetadataPostgreSQL + Prisma 7
🧭 VectorsQdrant (dense + sparse named vectors)
📐 Dense embeddingsBGE-small-en (fastembed)
🔤 Sparse embeddingsSPLADE++ EN v1 (fastembed)
🎯 RerankHugging Face cross-encoder (ms-marco-MiniLM-L6-v2)
📑 ParsingLlamaParse (LlamaCloud)
🤖 LLMOpenRouter
🌐 Web searchExa + LangGraph agent
🔭 ObservabilityLangfuse (OpenTelemetry)
--- ### 🏗️ Architecture ```text ┌────────────┐ presigned PUT ┌────────┐ │ Next.js │ ─────────────────────────────▶ │ MinIO │ │ web │ │(assets)│ └─────┬──────┘ └───┬────┘ │ │ │ REST (JWT) │ object key ▼ │ ┌─────┬──────┐ xAdd to files_stream ┌────▼─────┐ │ Express │ ───────────────────────────▶ │ Redis │ │ backend │ │ Streams │ └─────┬──────┘ └───┬──────┘ │ │ │ hybrid query + chat │ Dispatcher │ │ (routes by MIME) ▼ ▼ ┌───────────┐ ┌───────┬────────┐ │ Qdrant │ │ pdf_stream │ │ dense + │ ◀─── embed_stream │ image_stream │ │ splade │ │ audio_stream │ └───────────┘ embedding worker │ video_stream │ └───────┬────────┘ │ │ │ ┌───────────┐ │ └──────────▶ │ Postgres │ ◀──────────────┘ │ + users │ Parser workers │ + docs │ (per modality) │ + chunks │ │ + chats │ └───────────┘ ``` --- ### 🔄 Ingestion Pipeline Documents of any modality are accepted (PDF, images, audio, video). ```text ┌──────┐ presigned URL ┌───────┐ bytes ┌───────┐ │Client│ ─────────────────▶ │Express│ ────────────▶ │ MinIO │ └──┬───┘ └───┬───┘ └───────┘ │ │ │ POST /confirm │ ▼ ▼ ┌──────────┐ xAdd ┌──────────────┐ │ Document │ ─────▶ │ files_stream │ │ UPLOADED │ └──────┬───────┘ └──────────┘ │ Dispatcher │ (MIME detection) ┌─────────────────┼─────────────────┐ ▼ ▼ ▼ ┌────────────┐ ┌────────────┐ ┌────────────┐ │ pdf_stream │ │image_stream│ ... │video_stream│ └─────┬──────┘ └─────┬──────┘ └─────┬──────┘ │ │ │ ▼ ▼ ▼ ┌───────────┐ ┌───────────┐ ┌───────────┐ │ QUEUED │ │ PARSING │ │ PARSED │ └─────┬─────┘ └─────┬─────┘ └─────┬─────┘ │ │ │ ▼ ▼ ▼ ┌──────────────────────────────┐ ┌──────────────┐ │ ParsedChunkSet + Chunks │──▶ │ embed_stream │ │ (Postgres) │ └─────┬────────┘ └──────────────────────────────┘ │ ▼ ┌───────────┐ │ EMBEDDING │ └─────┬─────┘ │ ┌──────────┴────────────┐ ▼ ▼ ┌───────────┐ ┌───────────┐ │ Qdrant │ │ READY │ │ vectors │ │ (or FAIL) │ └───────────┘ └───────────┘ ``` > **Recovery**: Each stream has its own consumer group. Workers run a stale-job reclaimer (`XAUTOCLAIM`). After `MAX_RETRIES`, jobs move to a **Dead Letter Queue** (`dlq_stream`). --- ### 🔍 Retrieval & Chat ```text ┌──────────────────┐ │ User message │ └────────┬─────────┘ ▼ ┌──────────────────┐ │ Embed query │ │ dense BGE + │ │ sparse SPLADE │ └────────┬─────────┘ ▼ ┌──────────────────┐ │ Qdrant hybrid │ prefetch dense top-50 │ query │ prefetch sparse top-50 │ │ fuse with RRF → top 50 │ │ (filtered by user's docs) └────────┬─────────┘ ▼ ┌──────────────────┐ │ Cross-encoder │ │ rerank → top 5 │ └────────┬─────────┘ ▼ ┌──────────────────┐ │ System prompt + │ │ recent history + │ │ project prompt │ └────────┬─────────┘ ▼ ┌──────────────────┐ │ OpenRouter SSE │ │ stream │ └────────┬─────────┘ ▼ ┌──────────────────┐ │ Answer + sources │ │ stored on msg │ └──────────────────┘ ``` #### 💬 Chat Features | Feature | Details | |:--------|:--------| | 🔄 Streaming replies | Real-time SSE streaming from OpenRouter | | 📎 Source citations | Each answer carries ranked chunk references | | 📂 Projects | Organize chats with custom system prompts | | 📌 Pin / delete | Manage chat history | | ✏️ Edit & resend | Create version branches (1/2, 2/2) | | 🌐 Web mode | `/web` prefix triggers LangGraph research agent | | 📊 Live agent steps | Watch the web agent search, reason, and refine in real-time | | 📝 Conversation summaries | Rolling summaries injected into later prompts | --- ### 🔌 API Surface Base path: `/api/v1` (JWT middleware on all routes except `/auth/*`). | Area | Methods | Description | |:-----|:--------|:------------| | 🔐 **Auth** | `POST /auth/signup`, `POST /auth/signin` | User registration & login | | 📤 **Upload** | `POST /upload/post-file-url`, `POST /upload/confirm` | Presigned URL flow | | 📄 **Documents** | `GET /download/list`, `POST /download/get-download-url`, `DELETE /download/:id` | File management | | 💬 **Chat** | `GET /chat`, `GET /chat/:id`, `PATCH /chat/:id`, `DELETE /chat/:id`, `POST /chat/message` | Chat CRUD + SSE streaming | | 📂 **Projects** | `GET/POST /projects`, `PATCH/DELETE /projects/:id` | Project management | --- ### 🖥️ Frontend Routes | Route | Purpose | |:------|:--------| | `/` | Landing page (redirects to chat if signed in) | | `/signin`, `/signup` | Authentication | | `/dashboard` | Upload documents, view status, download / delete | | `/chat` | Full chat UI with history, projects, sources, web agent | --- ### 🗃️ Data Model (Postgres) | Model | Role | |:------|:-----| | `User` | Username + hashed password | | `Document` | Title, object key, `mimeType`, `modality`, status (`UPLOADED` → `READY` / `FAILED`) | | `ParsedChunkSet` | Group of parsed chunks per modality, status (`PARSED` / `INDEXED`) | | `ParsedChunk` | Individual text chunk with JSON metadata (page, timestamp, caption, OCR, etc.) | | `Project` | Named workspace + optional system prompt | | `Chat` | Title, pin, optional project, summary fields | | `Message` | role, content, `sourceChunks` JSON | | `Memory` | Schema for durable facts (not wired into chat yet) | > Chunk vectors live in **Qdrant** with payload including `documentId`, `chunkId`, `modality`, `page`, `timestamps`, `caption`. --- ### 💾 Storage Roles | Store | What it holds | |:------|:--------------| | 🗄️ **MinIO** | Original asset files (PDF, images, audio, video) | | 🐘 **PostgreSQL** | Users, docs, chunk sets, chunks, chats, messages, projects | | 📨 **Redis Streams** | Multi-stream job queue per modality + consumer group PEL + DLQ | | 🧭 **Qdrant** | Per-chunk dense + SPLADE vectors and text payload | > There is **no OpenSearch** — lexical signal comes from **SPLADE sparse vectors** inside Qdrant, fused with dense cosine via RRF. --- ### 🔭 Observability `@repo/langfuse` instruments: | Pipeline | Traced steps | |:---------|:-------------| | 💬 Chat RAG | `hybrid-retrieve` → `cross-encode-rerank` → `generate-response` | | 🌐 Web agent | LangGraph nodes (search → reason → refine → answer) | | 📄 Ingest | `process-document` and nested per-modality steps | If Langfuse keys are missing, tracing **no-ops** silently. --- **Key design decisions:** - 🪝 **Custom hook** (`useChatState`) encapsulates ~50 state variables and all SSE streaming logic - 🧩 **Presentational components** are pure — they receive props and render - 📦 **Types & helpers** are shared across all modules via local imports - 🔌 **Zero changes** to the route import (`import ChatPage from "@/components/chat-app"` resolves to `index.tsx`) --- ### 🏃 Local Development #### 📋 Prerequisites - [Bun](https://bun.sh) ≥ 1.3 - [PostgreSQL](https://postgresql.org) - [Redis](https://redis.io) - [MinIO](https://min.io) - [Qdrant](https://qdrant.tech) - API keys: LlamaCloud, OpenRouter; optional HF, Exa, Langfuse #### 🚀 Quick Start ```bash # 📦 Install dependencies bun install # ⚙️ Configure environment (root and/or apps/*) # Typical keys: DATABASE_URL, MinIO, JWT_SECRET, PORT, # STREAM_NAME / GROUP_NAME, COLLECTION, DENSE_DIM, # OPENROUTER_API_KEY, LLAMA_CLOUD_API_KEY, etc. # 🗃️ Apply Prisma migrations cd packages/db && bunx prisma migrate dev # 🏃 Start everything (web + backend + workers) bun run dev ``` #### 🎯 Run Individually ```bash bun run --filter web dev # 🖥️ Next.js on :3001 bun run --filter backend dev # ⚙️ Express on :3000 bun run --filter workers dev # 🏭 All workers bun run --filter workers dev:pdf # 📄 PDF worker only bun run --filter workers dev:embedder # 🧮 Embedder only bun run --filter workers dev:dlq # 🔄 DLQ worker only ``` #### 🔧 Development Commands | Command | Purpose | |:--------|:--------| | `bun install` | Install dependencies | | `bun run dev` | Start all apps via Turborepo | | `bun run build` | Production build | | `bun run lint` | ESLint (web: `--max-warnings 0`) | | `bun run check-types` | TypeScript check | | `bun run format` | Prettier (`--write`) | | `cd packages/db && bunx prisma migrate dev` | Apply migrations | --- ### 🧭 Design Principles | Principle | Description | |:----------|:------------| | ⚡ **Async ingest** | Uploads never block on parse/embed | | 🔍 **Hybrid retrieval** | Dense meaning + sparse terms, fused with RRF | | 📎 **Source grounded** | Answers carry chunk citations | | 🔒 **User scoped** | Retrieval and deletes are filtered by ownership | | 🧩 **Modular monorepo** | Shared clients in `packages/*` | | 🔭 **Observable** | Optional Langfuse traces end to end | | 🔌 **Decoupled parsing & embedding** | Re-embed without reparsing via `ParsedChunkSet` | | ➕ **Extensible modalities** | New types need only a new parser worker | ---
#### 🧠 RecallOS **Search your knowledge. Cite your sources.**
## SmolLM-135M_Med - **URL**: https://aadit032.github.io/portfolio/projects/2-smollm-135m_med/ - **Published**: 2026-02-01 - **Summary**: End-to-end pipeline for continued pretraining, supervised fine-tuning, and evaluation of SmolLM-135M on medical datasets. Medical-domain adaptation pipeline for `Aadit-032/SmolLM-135M_MedicalQA-SFT`. The project runs a two-stage training workflow — **continued pre-training (CPT)** on biomedical corpora, then **supervised fine-tuning (SFT)** on medical reasoning Q&A — with comprehensive evaluation at each stage. --- ## Overview | Stage | Base model | Data | Goal | Output | | --- | --- | --- | --- | --- | | **CPT** | `HuggingFaceTB/SmolLM-135M` | PubMed, PMC, Medline, FineWeb | Domain adaptation via next-token prediction | `Aadit-032/SmolLM-135M_Med-CPT` | | **SFT** | `Aadit-032/SmolLM-135M_Med-CPT` | medical-o1 CoT (default) or synthetic Q&A | Instruction following + medical reasoning | `Aadit-032/SmolLM-135M_MedicalQA-SFT` | --- Both stages use **LoRA adapters** via Unsloth with 4-bit (NF4) loading and 16-bit merged export. --- ## Pipelines ### CPT pipeline (`cpt.py`) ``` Load base model → baseline evals → CPT training → merge & save → post-training evals ``` 1. Load SmolLM-135M in 4-bit via Unsloth 2. Run baseline evals (perplexity, medical benchmarks, generation, lm-eval) 3. Download & tokenize 200K biomedical samples (`cpt_data.py`) 4. Train LoRA adapters for 1 epoch (`cpt_train.py`) 5. Merge and save to `SmolLM-135M_Med_Merged/` 6. Re-run all evals and compare before/after ### SFT pipeline (`sft.py`) ``` Load CPT model → pre-SFT evals → SFT training → merge & save → post-SFT evals ``` 1. Load `Aadit-032/SmolLM-135M_Med-CPT` (post-CPT checkpoint) 2. Run pre-SFT evals on the CPT model 3. Load SFT dataset from `sft_data.py` (medical-o1 by default) 4. Train LoRA adapters with response-only loss masking (`sft_train.py`) 5. Merge and save to `SmolLM-135M_Med-SFT-Merged/` (also hosted on Hugging Face) 6. Re-run all evals and compare pre-SFT vs post-SFT --- ## Usage ```bash # Install dependencies uv sync # --- CPT --- uv run cpt.py # full CPT pipeline uv run cpt_train.py # CPT training only uv run cpt_data.py # download & prepare CPT data only # --- SFT --- uv run sft_data.py # build/load SFT data (defaults to medical_o1) uv run sft.py # full SFT pipeline uv run sft_train.py # SFT training only ``` ### SFT data options ```bash # Default: medical-o1 reasoning dataset (cached after first run) uv run sft_data.py --loader medical_o1 # Synthetic: CPT split → semantic chunks → OpenRouter Q&A generation export OPENROUTER_API_KEY="your-key" uv run sft_data.py --loader synthetic --max-train-chunks 50 # Rebuild from scratch (ignore cached JSONL) uv run sft_data.py --rebuild # Choose output format: chat (default), alpaca, or raw uv run sft_data.py --format alpaca ``` --- ## Configuration All settings live in `config.yaml`: | Key | Value | Description | | --- | --- | --- | | `MODEL_NAME` | `HuggingFaceTB/SmolLM-135M` | Base model for CPT | | `SFT_MODEL_NAME` | `Aadit-032/SmolLM-135M_Med-CPT` | Starting checkpoint for SFT | | `SFT_LOADER` | `medical_o1` | SFT data loader (`medical_o1` or `synthetic`) | | `SFT_TEXT_FORMAT` | `chat` | Training text format (`chat`, `alpaca`, `raw`) | | `SFT_DATA_DIR` | `./data/sft/medical_o1` | Cached SFT dataset path | | `SEED` | `42` | Random seed | | `MAX_SEQ_LENGTH` | `512` | Max sequence length | | `train_file` | `./data/train.txt` | CPT training data | | `val_file` | `./data/val.txt` | CPT validation data | --- ## Training Details ### CPT (`cpt_train.py`) **LoRA** | Parameter | Value | | --- | --- | | Rank (`r`) | 32 | | LoRA alpha | 32 | | Target modules | `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj`, `embed_tokens`, `lm_head` | --- **Hyperparameters** | Parameter | Value | | --- | --- | | Epochs | 1 | | Per-device batch size | 32 | | Gradient accumulation | 4 | | Effective batch size | 128 | | Learning rate | 2e-5 | | Embedding LR | 2e-6 | | LR scheduler | Cosine | | Warmup ratio | 0.05 | | Packing | Enabled | --- **CPT data** (200,000 samples, 90/10 train/val split) | Source | Samples | Field | | --- | --- | --- | | PubMed Abstracts | 120,000 | `abstract` | | PMC | 40,000 | `text` | | Medline | 20,000 | `content` | | FineWeb | 20,000 | `text` | --- ### SFT (`sft_train.py`) **LoRA** | Parameter | Value | | --- | --- | | Rank (`r`) | 16 | | LoRA alpha | 16 | | Target modules | `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj` | --- **Hyperparameters** | Parameter | Value | | --- | --- | | Epochs | 1 | | Per-device batch size | 8 | | Gradient accumulation | 4 | | Effective batch size | 32 | | Learning rate | 2e-5 | | Loss masking | Response-only (`train_on_responses_only`) | | Packing | Enabled | --- **SFT data loaders** (`sft_data.py`) | Loader | Source | Description | | --- | --- | --- | | `medical_o1` (default) | `FreedomIntelligence/medical-o1-reasoning-SFT` | Medical CoT + answer pairs from DeepSeek-R1 | | `synthetic` | CPT train/val split + OpenRouter API | Semantic chunks → 4 Q&A pairs per chunk (factual, analytical, synthesis, unanswerable) | --- Both loaders append **20 general instruction pairs** (math, science, geography, refusal examples) to prevent catastrophic forgetting. SFT data is exported in three formats: **alpaca**, **raw** (`Question:` / `Answer:`), and **chat** (ChatML-style). --- ## Evaluation All results are saved to `./results/` as JSON. | Eval | File suffix | What it measures | | --- | --- | --- | | **Perplexity** | `_untrained`, `_trained`, `_pre_sft`, `_sft` | Sliding-window PPL on PubMed Abstracts & Medline (1,000 samples each) | | **Medical benchmarks** | same | PubMedQA (yes/no/maybe) & MedMCQA (4-option MCQ), 1,000 samples each | | **Generation** | same | 3 general + 3 medical prompts at varying temperature/top-k | | **LM-eval** | `results/lm_eval/*.json` | General capability: HellaSwag, PIQA, WinoGrande, ARC-Easy/Challenge, BoolQ | --- Medical benchmarks use **single-pass log-prob scoring** — one forward pass per question by batching all answer choices together. --- ## Project Structure ``` ├── cpt.py # CPT pipeline entry point ├── cpt_train.py # CPT LoRA training ├── cpt_data.py # CPT dataset download & preprocessing ├── sft.py # SFT pipeline entry point ├── sft_train.py # SFT LoRA training ├── sft_data.py # SFT dataset loaders (medical_o1 / synthetic) ├── model_utils.py # Model loading (base + SFT checkpoint) ├── config.yaml # All configuration ├── pyproject.toml # Dependencies ├── evals/ │ ├── benchmarks.py # PubMedQA & MedMCQA │ ├── perplexity.py # Sliding-window perplexity │ ├── generation.py # General + medical generation eval │ └── lm_eval.py # General capability benchmarks ├── data/ │ ├── train.txt # CPT training data │ ├── val.txt # CPT validation data │ └── sft/ # Cached SFT datasets (JSONL) └── results/ # Evaluation outputs ``` --- ## Dependencies - Python >= 3.13 - unsloth - datasets - omegaconf - evaluate - lm-eval Requires a CUDA GPU. For the synthetic SFT loader, set `OPENROUTER_API_KEY` in your environment. --- ## CPT Results | Metric | Untrained | Trained | Change | | --- | --- | --- | --- | | PubMed PPL | 18.76 | **15.03** | **-19.9%** | | Medline PPL | 14.24 | **11.39** | **-20.0%** | | PubMedQA | **49.5%** | 41.5% | -8.0 pts | | MedMCQA | 20.0% | **22.0%** | +2.0 pts | --- CPT achieved its primary objective — domain adaptation — with ~20% perplexity reduction on biomedical text. Downstream medical QA did not improve proportionally, motivating the SFT stage on medical reasoning data. --- ## SFT Results Evaluated on the CPT checkpoint (pre-SFT) vs the merged SFT model (post-SFT). Full outputs are in `./results/`. **Model:** `Aadit-032/SmolLM-135M_MedicalQA-SFT` | Metric | Pre-SFT (CPT) | Post-SFT | Change | | --- | --- | --- | --- | | PubMed PPL | 17.34 | **17.14** | **-1.2%** | | Medline PPL | 13.05 | **12.90** | **-1.2%** | | PubMedQA | 45.1% | **48.9%** | **+3.8 pts** | | MedMCQA | 24.2% | 24.1% | -0.1 pts | --- SFT recovered PubMedQA accuracy toward the untrained baseline (49.5%) while keeping biomedical perplexity stable. MedMCQA was essentially unchanged — a harder 4-option MCQ benchmark that likely needs more targeted training data or longer fine-tuning. **General capability (lm-eval, post-SFT)** | Task | Accuracy | | --- | --- | | HellaSwag | 34.5% | | PIQA | 68.2% | | WinoGrande | 51.6% | | ARC-Easy | 60.1% | | ARC-Challenge | 25.6% | | BoolQ | 59.8% | --- Generation samples (`generation_pre_sft.json` vs `generation_sft.json`) show modest gains in instruction-following structure, but outputs remain repetitive at 135M scale — expected for a model this size without RLHF or larger SFT corpora. ## Xcal - **URL**: https://aadit032.github.io/portfolio/projects/xcal/ - **Published**: 2026-02-01 - **Summary**: Realtime Excalidraw > A real-time collaborative whiteboard built with the HTML Canvas API, RoughJS, and WebSockets. Xcal is a browser-based whiteboard that enables multiple users to draw and collaborate on the same canvas in real time. Inspired by Excalidraw, it features hand-drawn rendering, low-latency synchronization, and a custom canvas engine built from scratch. --- ## Features - ✏️ Freehand drawing with smooth stroke interpolation - ⬜ Rectangle, Circle, Line, Arrow, and Text tools - 🎨 Hand-drawn rendering powered by RoughJS - 🔄 Real-time collaboration using WebSockets - 👥 Room-based collaborative editing - ⚡ Optimized rendering pipeline for smooth interactions - 🖱️ Shape selection, movement, and transformations --- ## Tech Stack | Layer | Technology | |--------|------------| | Frontend | React | | Rendering | HTML Canvas API | | Drawing Engine | RoughJS | | Backend | Node.js + Express | | Realtime | WebSockets | | Language | TypeScript | --- ## Architecture ```text User A User B │ │ └──────────┬───────┴ │ WebSocket Server │ Room-based Event Broadcast │ ┌─────────────────┴─────────────────┐ ▼ ▼ Canvas State Sync Drawing Events │ │ └──────────────┬────────────────────┘ ▼ Canvas Rendering Engine ``` --- ## Drawing Engine The canvas engine is implemented directly on top of the HTML Canvas API. Supported primitives: - Rectangle - Circle - Line - Arrow - Freehand Pencil Each shape is represented as structured data, allowing the canvas to be fully reconstructed on every redraw while supporting selection, movement, and editing. --- ## Rendering Pipeline To maintain smooth interactions, the rendering system: - Redraws only when necessary - Minimizes expensive canvas operations - Batches user interactions - Efficiently replays canvas state This allows the application to maintain approximately **45 FPS** during collaborative drawing sessions. --- # Blog / Articles 1 article — full text below. ## Building a 25M Parameter GPT from Scratch - **URL**: https://aadit032.github.io/portfolio/blog/litegpt-25m/ - **Published**: 2026-07-04 - **Summary**: Architecture choices, tokenizer mistakes, training issues, and lessons from training LiteGPT on FineWeb and TinyStories. "How hard can this be?" To satisfy my curiosity and understand every design choice firsthand, I built a 25M parameter GPT from scratch. This article is a collection of the lessons I learned along the way. ### TLDR I trained a 25M parameter language model from scratch on 500M tokens from FineWeb and TinyStories, using modern architecture choices such as RoPE, GQA, SwiGLU, and RMSNorm. The final run used an NVIDIA RTX A5000 GPU for roughly two hours. Model weights and logs are on [Hugging Face](https://huggingface.co/Aadit-032/LiteGPT-25M). The training code is on [GitHub](https://github.com/Aadit032/Lite-GPT). ### Why bother building this? Building a model from scratch taught me how much the pre-training data matters, gave me a more concrete intuition for scaling laws, and helped demystify a system that previously felt like a black box. I was inspired by Karpathy's NanoGPT. It made me think about how I could improve the architecture, try different dataset mixtures, and push for better results with fewer parameters and less compute. I wanted the repository to be easy for beginners to read while still being flexible enough for future experiments. ### Section 1: Architecture The earlier version [LiteGPT-16M](https://huggingface.co/Aadit-032/LiteGPT-16M) followed a GPT-2 style architecture with LayerNorm, multi-head attention, GELU feed-forward layers, and learned positional embeddings. For the 25M model, I modernized the architecture with RMSNorm, SwiGLU, RoPE, and grouped-query attention. These choices make training faster and scale better. ![Transformer architecture](https://pbs.twimg.com/media/HMZOmxLawAA77Z_.jpg) #### 1.1) Dataset I started with a dataset mixture containing **60% Fineweb, 30% TinyStories and 10% The Stack Smol** from HuggingFace because I wanted to fine-tune the base model later to create a small coding agent for fun but I had to **remove the code dataset** completely as I wanted the model to understand natural language better and the coding data introduced a very different token distribution than natural language and I thought that it would help the model **generalize** better. The dataset I finally settled on and where the model was finally trained on was **60% FineWeb and 40% TinyStories**. I expected the loss to go down by a lot after I removed the The Stack Smol dataset keeping only natural language data but this didn’t change the loss by too much. >Next experiment would be to train this model exclusively on FineWeb or TinyStories. I am quite sure I can get the model’s validation loss to go down. #### Tokenizer I initially used the GPT-2 tokenizer, the same one used in NanoGPT. I could not get validation loss below 4, which suggested something was fundamentally wrong. The issue was vocabulary size. GPT-2's vocabulary has 50,257 tokens. With `d_model = 320`, the embedding table alone used roughly **16M parameters out of the 25M** total trainable parameters. That left too little capacity for the transformer itself. To fix this, I trained a byte-level BPE tokenizer with a **vocabulary size of about 16K** using Hugging Face Tokenizers. The dataset is stored as `uint16` binary files and loaded lazily with `np.memmap()`, which lets the data loader grab random chunks without loading the whole dataset into RAM. >500M tokens at uint16 is 500,000,000 x 2 bytes ~ 1GB >np.memmap() allows loading files bigger than the RAM and has a low startup time because it only loads the pages that are being read, directly from the disk. #### 1.2) Scaling laws >Before **Chinchilla**, models like GPT-3 were undertrained (175B params trained on ~300B tokens). Chinchilla scaling laws suggest that for a fixed compute budget, model size and number of training tokens need to be balanced. The rough compute-optimal regime is about **20 training tokens per parameter**. For a 25M parameter model, that points toward roughly **500M training tokens**. This is why I targeted a 500M-token dataset. #### 1.3) Learning rate schedule The learning rate linearly warms up to a peak of `6e-4` and then follows cosine decay. Warmup matters because early losses are large, and using a high learning rate immediately can make updates unstable. Cosine decay helps reduce update size as the model starts converging. ![lr_graph](https://pbs.twimg.com/media/HMZN_m0bwAAw8Ff.png) #### 1.4) Transformer block #### 1.4.1) RMSNorm The model uses a GPT-2 style **pre-norm** block, but with RMSNorm instead of LayerNorm. RMSNorm normalizes by root-mean-square magnitude and avoids subtracting the mean, which makes it cheaper than LayerNorm. #### 1.4.2) RoPE RoPE rotates the query and key vectors before attention, allowing the attention mechanism to model **relative positions**. I chose RoPE because it reduces learned positional parameters and tends to **generalize better** to longer contexts than learned positional embeddings. ![rope](https://amaarora.github.io/images/rope-implementation.png) #### 1.4.3) GQA The attention block implements **Grouped Query Attention (GQA)**, where there are n_heads query heads but only n_kv_heads key and value heads. Multiple query heads share the same key and value heads, significantly **reducing the memory and computation required** for the KV cache during inference. Different attention heads learn different attention patterns. Empirical studies have shown that sharing keys and values across multiple query heads has little impact on model quality, making GQA an effective trade-off between efficiency and performance. ![gqa](https://pbs.twimg.com/media/HMZPD7KakAArHtI?format=png&name=small) Attention is computed using `torch.nn.functional.scaled_dot_product_attention()`, which dispatches to highly optimized **fused kernels** that combine the attention operations into a single implementation, **reducing memory accesses** and improving throughput. When the hardware and input satisfy the required conditions, PyTorch automatically uses FlashAttention, providing further speedups and lower memory usage without requiring changes to the model code. #### 1.4.4) Residual add After the attention block, the original input is added back to the output of the attention layer through a residual connection. Residual connections **preserve** the original representation while providing an path for gradients, making optimization significantly easier. They provide a direct path for gradients to flow during backpropagation, helping avoid the vanishing gradient problem as models become deeper. This allows each transformer block to learn a **small refinement** to the input instead of having to learn an entirely new representation from scratch, resulting in faster convergence and more stable training. The same residual connection is also applied after the SwiGLU feed-forward network, making every transformer block responsible for learning only incremental improvements to the representation. #### 1.4.5) SwiGLU MLP/FFN Then comes the SwiGLU MLP, here, instead of one up-projection as done in GELU, there are **2 up-projections**, one is the original up-projection and the other is the gate projection. This gating feature decides which features are worth keeping and which aren’t and **suppresses** those features thus acting as a **feature gate**. #### 1.4.6) Cross entropy loss After going through n_layers of transformer blocks, it goes through a final RMSNorm and produces the final logits which are then passed through softmax to return **probabilities over n_vocab choices**. The loss is calculated here using cross entropy loss which takes the **negative log probability** assigned to the correct next token. The negative log is taken because it strongly penalizes when the model is confidently wrong. The farther the probability is from the correct answer, the stronger the penalty is. #### 1.4.7) Generation The model doesn’t output probabilities by default, when the model is being used for inference, these **logits are divided by the temperature** and passed through a **softmax function** that gives us the probabilities of the next token over the entire vocabulary of the model. Temperature controls the **sharpness of the output probability distribution**. Higher temperature produces a flatter distribution with more similar probabilities, while lower temperature produces a sharper distribution where the highest-probability tokens dominate. ### Section 2: Implementation This is the configuration of the model: | Parameter | Value | | --- | --- | | n_layers | 8 | | batch_size | 64 | | grad_accum_steps | 2 | | d_model | 448 | | n_heads | 8 | | head_dim | 56 | | n_kv_heads | 4 | | ffn_dim | 1152 | | context_length | 512 | | vocab_size | 16384 | --- The hidden dimension follows the LLaMA-style rule of roughly `8 / 3 * d_model` instead of the classic GPT-2 `4 * d_model`, because SwiGLU uses three linear projections rather than two. Parameter count: | Component | Params | | --- | --- | | Token Embedding | 16384 × 448 = 7,340,032 | | Attention (GQA) | [(448×448) + (448×224) + (448×224) + (448×448)] × 8 = 4,816,896 | | SwiGLU FFN | [(448×1152) + (448×1152) + (1152×448)] × 8 = 12,386,304 | | RMSNorm | (448 × 2) × 8 = 7,168 | | Final RMSNorm | 448 | | LM Head | weight tied with token embeddings | | Total | **≈ 24.6M** | --- Training ran for 40,000 forward and backward iterations with a gradient accumulation factor of 2, resulting in 20,000 optimizer steps. The total number of tokens seen during training was: >64 x 512 x 40,000 = 1,310,720,000 tokens That is about 1.3B tokens, or roughly 2.6 epochs over the 500M-token dataset. ### Section 3: Issues I faced #### 3.1) Hardware constraints I started on the free-tier T4 GPU on Google Colab, but it was too slow and sessions were terminated aggressively. The **16GB VRAM limit** also forced a smaller batch size. The memory used by a model in training is roughly: >Weights + Gradients + Optimizer states + Activation I switched to RunPod and rented an **NVIDIA RTX A5000**, which let me increase batch size, hidden dimension, and other hyperparameters. There are 2 ways to reduce the VRAM used on GPUs: 1. **Gradient checkpointing**: Instead of storing all the activations of the model, store only a few activations and when backpropagating, recompute the activations. 2. **Gradient accumulation**: This increases the effective batch size by accumulating gradients over multiple forward/backward passes before performing one optimizer update. #### 3.2) Loss not going down The loss was not budging and the model was not able to get to a good loss value, which meant that something in my setup was broken. After spending quite some time looking around I realized that my tokenizer had n_vocab of 50257 and the embedding layer ended up taking 16M params out of the 25M trainable ones and that’s why the model was just not able to generalize all that data when the transformer only had 9M params to train on. #### 3.3) Steps are not the right training metric I initially thought training for more steps would automatically improve the model. That was wrong. The **number of quality tokens seen during training** is a much better thing to focus on. For LiteGPT-16M, the training loss was around 2 while validation loss was around 5. That was a clear overfitting signal: the model memorized the training data but **failed to generalize**. #### 3.4) Wrong data mixture The initial mix included coding data, but the model was too small to produce useful code. The code dataset also **pulled the token distribution away** from the natural language behavior I cared about most. ### Section 4: Improvements to make 1. Train exclusively on FineWeb to test a larger and more diverse natural language distribution. 2. Scale to more than 1B tokens while following Chinchilla-style scaling laws. 3. Fine-tune on instruction data for a specific downstream task. 4. Try alternative optimizers such as AdamW-mini or Sophia. ### Section 5: Example generations #### Prompt: The capital of France is > well known for its great geographical experience. The historical setting is not one of a sort of the greatest of all the times, but it's a central part of the history of the country. It's a historical town that has seen a huge shift inside this famous city of the same name. The city is known as "The Dolceus" and also known as "The Dolceus." It's a place where a lot of people live and live. #### Prompt: Once upon a time, there lived > a little girl named Lily. She was very excited because she was going on a trip to see Grandma. She helped her mom pack their luggage with clothes and toys. Then, they got in the car and drove to Grandma's house. As they drove, Lily saw something strange. It was a big, gray cloud that seemed to float away. She pointed and asked her mom what it was. Her mom said it was a mild cloud, but it was a very hot day. #### Prompt: The dragon opened its eyes and > looked around. It saw something shiny and colorful. It jumped and grabbed the object with its sharp claws. The dragon flew away and the dragon breathed a sigh of joy. It flew back up into the sky and found its owner, who was so proud of it. He said the dragon had been an amazing hunter. #### Prompt: Tom had a little red ball. > Lily had a red ball and Tom had a blue ball. They both loved to play with their toys. But they both wanted the red ball. They both pulled it from each other. "Give me the red ball!" Lily shouted. "It's mine!" "No, it's mine!" Tom yelled. "It's mine!" They pulled so hard that they separate the red ball from each other. They both fell down and started to cry. Mom heard the noise and came to see what was wrong. She saw the red ball on the floor and the tears on Lily's face. She was not angry, but she was sad. "Why are you fighting?" Mom asked. "You both have the red ball, but you both have the same toy." "That is boring!" Lily said. "You both need to share and take turns." "Okay!" Mom said. "But you can both have the red ball. ### Section 6: Results #### Final metrics | Metric | Value | | --- | --- | | Train Loss | 2.685 | | Val Loss | 2.825 | | Perplexity | 16.86 | --- #### GPU Benchmark (Inference on A5000) | Metric | Value | | --- | --- | | Throughput | ~75K tok/s | | Latency (mean) | ~4.2 ms | | Latency (P95) | ~5.8 ms | | GPU Memory (alloc) | ~0.9 GB | ### Section 7: Observations The model performs best on story-like prompts, which is not surprising given the TinyStories mixture. Factual accuracy is poor at 25M parameters. Repetition loops become common at higher temperatures. Coding generations are essentially non-functional. Reasoning is limited, but the model does show a basic grasp of narrative structure. ### Section 8: References 1. [Attention Is All You Need, Vaswani et al.](https://arxiv.org/abs/1706.03762) 2. [NanoGPT, Karpathy](https://github.com/karpathy/nanoGPT) 3. [Language Models are Unsupervised Multitask Learners, Radford et al.](https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf) 4. [LLaMA: Open and Efficient Foundation Language Models, Touvron et al.](https://arxiv.org/abs/2302.13971) 5. [RoFormer: Enhanced Transformer with Rotary Position Embedding, Su et al.](https://arxiv.org/abs/2104.09864) 6. [GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints, Ainslie et al.](https://arxiv.org/abs/2305.13245) 7. [Scaling Laws for Neural Language Models, Kaplan et al.](https://arxiv.org/abs/2001.08361) 8. [Training Compute-Optimal Large Language Models, Hoffmann et al.](https://arxiv.org/abs/2203.15556) 9. [FlashAttention, Dao et al.](https://arxiv.org/abs/2205.14135) 10. [RMSNorm, Zhang and Sennrich](https://arxiv.org/abs/1910.07467) --- _End of agent context._