Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Anecdotal data warning but for context my research is in medical informatics and I've quite extensively followed publications on transformers dating back to the early BERT variants (including non-medical).

I'm making that statement as my experience (easily several hundreds of publications read or reviewed over 3 years) is that it is very uncommon to see TPU's mentioned or TRC acknowledged in any non-Google transformer paper (especially major publications) dating back to the early BERT family of models despite the fact that Google is very generous with research credits (they'll give out preemptible v3-32s and v3-64s for 14 days with little question, presumably upgraded now as I haven't asked for credits in a while).

Fully acknowledge this isn't quality evidence to back my claim and I'm happy to be proven wrong but I'm very confident a literature review would support this as when I tried to use TPUs myself I couldn't find much.

This doesn't account for industry use, there is probably a non-insignificant amount of enterprise customers still using AutoML (I can think of a few at least) which I believe uses the TPU cloud but I would be surprised if many use TPU nodes directly outside of Jax shops like cohere and anyone still using TF.

PyTorch XLA has just breaks too much otherwise and when I last tried to use it in January of this year there was still quite a significant throughput reduction on TPUs. Additionally when using nodes there is a steeper learning curve on the ops side (VM, storage, Stackdriver logging) that make working with them harder than spinning up a A100x8 which is relatively cheap, cheaper than the GCP learning curve for sure.



> Anecdotal data warning but for context my research is in medical informatics

Isn't Medical Informatics inherently biased against the cloud? That's my uninformed guess as an outsider.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: