I was roughly using estimates for as single Nvidia GB300 NVL2 superchip, which has about 1.3 TB of memory (which I think is probably ballpark the size of the distilled models that are being used for inference).
Timing-wise, I'm making an extremely liberal estimate. (In reality, Sora 2 takes about 2 minutes to generate a video in my testing).
The thing that matters here is ultimately the product of time and memory (the latter of which I use as a stand-in for calculating number of chips being used), so I'm more or less asserting that the models being used to generate the free Sora 2 videos have less than 6 trillion parameters (since 2 minutes being 1/5 of 10 minutes give us a 5x multiplier we could apply to 1.3 TB), or less than 24 trillion parameters with aggressive quantization. I can explain why I think parameter count is less than 6 trillion if you'd like, but it's pretty hand-wavy intuition.
Timing-wise, I'm making an extremely liberal estimate. (In reality, Sora 2 takes about 2 minutes to generate a video in my testing).
The thing that matters here is ultimately the product of time and memory (the latter of which I use as a stand-in for calculating number of chips being used), so I'm more or less asserting that the models being used to generate the free Sora 2 videos have less than 6 trillion parameters (since 2 minutes being 1/5 of 10 minutes give us a 5x multiplier we could apply to 1.3 TB), or less than 24 trillion parameters with aggressive quantization. I can explain why I think parameter count is less than 6 trillion if you'd like, but it's pretty hand-wavy intuition.