Tweet by madiator
September 20, 2025
I love how he has articulated what he wants clearly. But people still don't seem to understand what this is. It's not RAG, it's not notebooklm, it's not lmstudio. It's not a solved problem at all, contrary to what anyone claims. I have also always wanted to something like this, but never got to it due to technical hurdles, but also possibly the business angle. But let's aside the business angle for a minute. Please note that I led a Neurips paper called "Generative Retrieval" (https://t.co/qOT30po0RX) where we transfer the "recommendation knowledge" into a transformer. So that's where my interest in this field comes from. A reasonable start is to do CPT (continual pretraining) of the unstructured text (but needs some data curation) followed by some sort of SFT (and maybe even RL), all of which can be curated. The paper to read here, for the former, is "Synthetic Continual pre-training" from Stanford: https://t.co/D8sPyGNTuy (besides the above one). And it's worth pointing out that ultimately this is first and foremost a data curation problem, and then a modeling problem. The simplest form of what he is saying is to take a pdf (something that's reasonably long) and then train a llm out of it such that it has internalized the knowledge in the pdf and no RAG is needed. Again, this seems easy but it's not. And then of course there is a lot of fascinating technical problems here, besides the challenges around cost (CPT itself is expensive, but the data gen process -- if you read the above paper -- is also expensive). For example, CPT can make the model a bit dumber in other aspects. So how do you methodically add existing pretraining data (which people don't have access to) during the CPT process.
- Author
- madiator
- Date
- September 20, 2025
- Canonical URL
- /tweets/madiator-1969477887943459226-fff32c