Patel explains long context windows fail because pre-training uses short context and lacks data for longer ranges.
“But it's like, what data do I have? That's why my two fifty ks to one mil context is trash anyways, because there's no data on this stuff.”
post Patel says OpenAI explicitly stated at Hot Chips they bet against disaggregation
https://x.com/dylan522p/status/2092633971041403333
post Patel claims evidence exists of organizations astroturfing false datacenter narratives with non-domestic funding
https://x.com/dylan522p/status/2091946535022215304
post Patel says SemiAnalysis tracks changes as they happen, with majority revenue from industry not hedge funds
https://x.com/dylan522p/status/2088653662780547546
post Patel says SemiAnalysis has more China compute mapped than cited estimate
https://x.com/dylan522p/status/2084461665173848493
post Patel says datacenter model hard to justify selling to individuals, only B2B makes sense for effort required
https://x.com/dylan522p/status/2082365159750730226
post Patel asks Douglas if he can use the model without privacy protections
https://x.com/dylan522p/status/2092262188409119109