A Jev-like wrapper for LLMs, including vision models

I was intrigued by Jev and the self-hostable projects appearing around it, such as OpenJev and SemIf. Reading about them introduced me to a neat trick: reading an LLM's token probabilities.

Apparently this is an old trick for some people. See e.g. OpenAI's logprobs cookbook. But it was new to me.

I believe the basic idea is to write a prompt like this:

State: My order arrived broken and I want a refund. Question: Which team should handle this? [A] billing [B] shipping [C] returns Answer with the letter of the best option only.

Then add a few JSON request parameters to a compatible Chat Completions request: