Model Size Benchmark and the Efficiency Leap from Compiled Languages
Confirms that compiled languages like TileLang bring efficiency gains, not losses, with hardware execution efficiency loss of only 1% to 2%.
Yes, it's a huge efficiency gain. So this is an opportunity.
Second question, you mentioned earlier that we're now using some compiled languages other than CUDA for inference. As I understand, we used to rely on the NVIDIA ecosystem, probably using PTX a lot. Now that we're using more of the TileLang you just mentioned, will that significantly reduce our inference efficiency, or at least in the short term? I'm not sure how you view the efficiency loss from this shift to compiled languages, or whether in the long run it's actually a supplement that improves efficiency?
It improves efficiency, significantly improves efficiency.
OK, so no negative impact?
Yes, it's a huge efficiency gain. So this is an opportunity. It's like before, you couldn't leave the CUDA ecosystem, but now we can abandon it and use a simpler approach with TileLang. It's a high-level language, writing programs is fast, and the codebase is small; we can rewrite it.
I see. So both of these are actually, as you mentioned, a big opportunity brought by AI, not a shortcoming that needs to be addressed.
Yes, it's a big opportunity for technological development. It's not about AI opportunity, because we also have a project where we're using AI to write TileLang.
Would that be even faster?
Right now, all TileLang is written by humans, but it's already much faster than writing CUDA was before.
I see. So even for low-level hardware execution efficiency, you think there's no impact?
A loss of 1% to 2%, I think that's acceptable.
OK, understood. Alright, thank you, that's very clear.
Anyone else have questions? No? Let's wrap up for today. OK, thank you. Thank you all very much for your time. Thank you, Mr. Liang. Thank you. Bye.
Bye.