Which one is better?
It depends on the metric, so here is every metric we measure, and who wins each. The numbers come from the same runs as the Results.
What to keep in mind
- Ten episodes per game. Enough to see large differences, not enough to split close ones. The Results page shows the ranges.
- Decision time is not a fair race yet. Jev answers over the network with rate limits; Laya ran on a laptop CPU, not the GPU it is built for, where its makers report about 33 ms.
- These games are not what Laya was trained for. It was trained on support tickets, moderation and similar work. Its weights are open, so it can be fine-tuned on decisions like these; that is the next experiment.