Everything else is fine, but the requests are too slow and there are too few models.
ds+zcode was actually pretty impressive to use. I made two animations with remotion, and for a pure text model, the results weren’t bad. But there were a lot of visual overlaps and ugly bits that had to be pointed out by hand, and it could only generate fairly simple vector graphic animations. Anything more complex would break down completely, and it can’t crawl web assets on its own—I had to fetch them myself (even if it could crawl, it wouldn’t understand them). Hooked up skill to mimo2.5 as its “eyes,” but the results weren’t great either. 【Science Animation】Timestamps and the 2038 Problem That Even a Paramecium Could Understand
Not sure if it’s worth switching to a multimodal model. Multimodality, coding ability, and cost have practically become an impossible triangle now
I’ll optimize MODELOC and similar projects when things slow down; also, I’m discussing GPU sharing with some organizations, and I might deploy open-source models like KIMI, GLM, and DeekSeek myself for sharing.