Rendered at 19:53:32 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
dlcarrier 20 hours ago [-]
Can you also measure how often the LLM response makes people laugh? Sometimes the responses that aren't attempting a joke are the funniest, and I'd be more interested in stats of which LLM succeed in that metric.
killiandunne1 19 hours ago [-]
Interesting point - thoughts on how to do this? Honestly most models are not-to-kinda funny so I'd be surprised if there were many lol moments. Oral delivery is something I think is v interesting though
dlcarrier 3 minutes ago [-]
You'll probably still get plenty of smirks, smiles, and silent chuckles. OpenCV or a similar machine vision system should be able to pick those out from a webcam aimed at a reviewer.
maxzhdev 15 hours ago [-]
I can’t imagine how to correctly assess the ability to humor, because this is a very subjective and relative phenomenon. The work ahead is serious
killiandunne1 2 hours ago [-]
We don't wanna be too serious ;)
rafaepta 20 hours ago [-]
measuring humor might be halfway to measuring taste. congrats on this... really original contribution. wonder if you're planning to evolve the benchmark to incorporate a multi-language dimension. would love to see how Mistral and models built outside the US would perform.
killiandunne1 20 hours ago [-]
Nice a good way to 10x inference costs... worth seeing though ahaha