Thanniru Sai Teja
AI Systems Researcher and Full-Stack Engineer. 4th-year B.Tech CSE (Data Science) student at BVRIT Narsapur, Telangana, India (2023-2027, CGPA 8.85/10). Builds attention-free sequence architectures and memory systems on consumer hardware, and ships performance patches upstream to curl, nvm, colibri, llmfit and Soup. Open to SWE and ML roles.
Open source
- llmfit: 46% faster model search. Short-circuiting field checks in the TUI's apply_filters(), replacing ~2,000 heap allocations per keypress.
- nvm: 141x faster nvm_tree_contains_path. Replaced repeated dirname calls with an in-memory POSIX parent walk. ~350 process forks down to 0.
- Soup: 30% faster dataset validation. A single-pass validator with hashable row signatures and no intermediate allocations.
- curl: 36% faster Base64 decode. Rewrote the decode lookup table and quantum loop. Benchmarked by maintainer Daniel Stenberg over 10M operations and committed upstream.
- colibri: 99.6% less allocation overhead. Swapped 12 malloc/free calls per batch in the DeepSeek V4 indexer for a persistent 32-byte-aligned scratch arena.
- Windows-MCP: 52% faster coordinate resolution. Replaced an N+1 lookup in MultiEdit and MultiSelect with bulk resolution: 37.45% faster at 3,000 resolutions, 52.09% at 35,000. My first open-source contribution.
Projects
- HGDM: 100% attention-free byte-level sequence model. 1.006B parameters pushed to 1.46B training tokens on a 4GB VRAM GPU. O(1) inference memory across 20x context scaling.
- NCM: Tensor-based episodic memory with four-dimensional retrieval: semantic, emotional, state-conditioned, and temporal. The same query can recall different memories depending on the system's state.
- Small-model RAG study: 60,000 controlled evaluations on 600 HotpotQA questions across five RAG architectures. The 350M model beat the 700M one on Token F1 (11.07% vs 7.27% with Naive RAG).
- ZS-ISAB: Scales TabPFN's attention from quadratic to linear with seeded anchor selection, an online softmax accumulator, and a zero-shot attention mask, enabling 500K+ row inference on 4GB VRAM.
- NOUS: A bare-metal operating system for x86_64 where memory is addressed by semantic embeddings rather than pointers, and the resident kernel is a recurrent neural model.
- Compute Pool: A Rust CLI with a Python training runtime that pools up to 8 Kaggle accounts into one cluster: independent jobs, data-parallel training with an SSH all-reduce, and pipeline-parallel training.
- Event Ticket Booking System: End-to-end ticketing platform with real-time booking, QR code ticket generation, PDF export, and a Cloud Functions backend. Served 300+ real users across campus events.
- STUD Recruitment Platform: Built and shipped a production recruitment platform for Stud Entertainments in 3 days flat, with admin tracking, automated interview scheduling, and auth.
- Orbius AI: Built at the NxtWave × OpenAI Academy state-level buildathon. Multi-Agent Mode lets GPT, Gemini, Perplexity and other models work on the same problem together.
- FormatForge: Upload one photo, get platform-ready assets for Amazon, Flipkart, Instagram, Spotify with natural language in-place AI re-editing.
- Chaff App: Web-based farmer-to-manufacturer stubble marketplace turning agricultural waste into valuable industrial resources.
Writing · llms.txt · GitHub · LinkedIn