{"article":{"slug":"a-jev-like-wrapper-for-llms-including-vision-models","title":"A Jev-like wrapper for LLMs, including vision models","subtitle":null,"summary":"Allan shows a small single-function Jev-style wrapper for LLMs that also handles vision models, with practical code for local and API backends.","content_type":"tutorial","language":"en","canonical_url":"http://allanrbo.blogspot.com/2026/09/a-jev-like-wrapper-for-llms-including.html","author":{"name":"Allan R. Bo","url":"http://allanrbo.blogspot.com/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Allan's Blog","url":"http://allanrbo.blogspot.com/","listing_slug":null,"listing":null},"topics":[{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"Tutorial","slug":"tutorial","url":"https://listedarticles.com/topics/tutorial"},{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1575,"reading_minutes":7,"published_at":"2026-09-26T12:00:00.000Z","added_at":"2026-09-26T06:12:53.038Z","updated_at":"2026-09-26T06:12:53.038Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/a-jev-like-wrapper-for-llms-including-vision-models","markdown_url":"https://listedarticles.com/articles/a-jev-like-wrapper-for-llms-including-vision-models.md","example":false,"citation":"Allan R. Bo, Allan's Blog. \"A Jev-like wrapper for LLMs, including vision models.\" 26 Sept 2026. http://allanrbo.blogspot.com/2026/09/a-jev-like-wrapper-for-llms-including.html (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"http://allanrbo.blogspot.com/2026/09/a-jev-like-wrapper-for-llms-including.html"},"body_markdown":"# A Jev-like wrapper for LLMs, including vision models\n\nI was intrigued by [Jev](https://docs.typesafe.ai/introduction) and the self-hostable projects appearing around it, such as [OpenJev](https://huggingface.co/openjev/openjev) and [SemIf](https://github.com/TheoLeeCJ/SemIf-OpenJev). Reading about them introduced me to a neat trick: reading an LLM's token probabilities.\n\nApparently this is an old trick for some people. See e.g. [OpenAI's logprobs cookbook](https://developers.openai.com/cookbook/examples/using_logprobs). But it was new to me.\n\nI believe the basic idea is to write a prompt like this:\n\nState: My order arrived broken and I want a refund.\nQuestion: Which team should handle this?\n[A] billing\n[B] shipping\n[C] returns\nAnswer with the letter of the best option only.\n\nThen add a few JSON request parameters to a compatible Chat Completions request:\n\n```\n{\n  \"max_completion_tokens\": 1,\n  \"logprobs\": true,\n  \"top_logprobs\": 20\n}\n```\nThe LLM API will return the letter plus the model's log probabilities for alternative tokens.\n\nRepeat for each question. Forcing it to generating only one token avoids a lengthy answer and is super quick, though processing the input still costs time. Though for each of the questions a shared state prefix can be KV-cached if the backend supports it.\n\nThe fun part: this works with vision models too. Jev's [documented request format](https://docs.typesafe.ai/introduction/quickstart) currently describes only text/JSON state. I added an `attachments` field for images for my local experiments.\n\nMy example captures webcam frames, sends base64 JPEGs, and prints a table: is a person visible, are we indoors or outdoors, and how bright is the scene? With Gemma 4 12B on my RTX 3090, I get around **1 frames per second**, with three questions per frame. I also ran it against OpenAI gpt-6-luna and got around 0.2 FPS. Presumably because I didn't make any effort to avoid the cost of a separate connection through their system per question per frame.\n\nSpecialized computer vision models surely are much more efficient, but what I like here is the flexibility: change a condition by describing it in plain text.\n\nHere's the standalone Python example (OpenCV is just used for convenient access to the webcam, not for any actual computer vision):\n\n```\n#!/usr/bin/env -S uv run --script\n# /// script\n# dependencies = [\"opencv-python\"]\n# ///\n\"\"\"Preview and score webcam frames with llama.cpp or OpenAI.\nuv run webcam.py\nuv run webcam.py https://api.openai.com/v1 gpt-6-luna\nOpenAI reads OPENAI_API_KEY.\n\"\"\"\nimport argparse\nimport base64\nimport concurrent.futures\nimport datetime\nimport json\nimport math\nimport mimetypes\nimport os\nimport pathlib\nimport time\nimport urllib.parse\nimport urllib.request\nimport cv2\n# attachments is our custom addition to the Jev request format.\ndata = json.loads(\"\"\"\n{\n    \"state\": \"Inspect this webcam frame. Judge only what is visibly present.\",\n    \"attachments\": [],\n    \"questions\": {\n        \"person\": {\n            \"type\": \"noul\",\n            \"instructions\": \"Is a person visible?\"\n        },\n        \"plant\": {\n            \"type\": \"noul\",\n            \"instructions\": \"Is a plant visible?\"\n        },\n        \"setting\": {\n            \"type\": \"choice\",\n            \"instructions\": \"Where is the camera?\",\n            \"criteria\": {\n                \"indoors\": null,\n                \"outdoors\": null,\n                \"unclear\": null\n            }\n        },\n        \"light\": {\n            \"type\": \"score\",\n            \"instructions\": \"How bright is the scene?\",\n            \"criteria\": [\n                \"dark\",\n                \"dim\",\n                \"bright\"\n            ]\n        }\n    }\n}\n\"\"\")\ndef score(data, url, model):\n    state = data[\"state\"]\n    if not isinstance(state, str):\n        state = json.dumps(state)\n    # Attachments are our extension to the Jev-style request format:\n    # image file paths or base64 data URLs. Load them once for all questions.\n    images = []\n    for attachment in data.get(\"attachments\", []):\n        if attachment.startswith(\"data:image/\"):\n            images.append(attachment)\n            continue\n        path = pathlib.Path(attachment).expanduser()\n        mime_type, _ = mimetypes.guess_type(path)\n        if mime_type not in {\"image/png\", \"image/jpeg\", \"image/webp\", \"image/gif\"}:\n            raise ValueError(f\"Unsupported image file: {path}\")\n        encoded = base64.b64encode(path.read_bytes()).decode()\n        images.append(f\"data:{mime_type};base64,{encoded}\")\n    # Send the API key only to OpenAI.\n    is_openai = urllib.parse.urlsplit(url).hostname == \"api.openai.com\"\n    headers = {\"Content-Type\": \"application/json\"}\n    if is_openai:\n        headers[\"Authorization\"] = \"Bearer \" + os.environ[\"OPENAI_API_KEY\"]\n    answers = {}\n    for name, question in data[\"questions\"].items():\n        # Represent choices, booleans, and ordinal levels as lettered options.\n        if question[\"type\"] == \"choice\":\n            options = question[\"criteria\"]\n        elif question[\"type\"] == \"noul\":\n            options = {\"true\": None, \"false\": None} | question.get(\"criteria\", {})\n        elif question[\"type\"] == \"score\":\n            options = {str(i): description for i, description in enumerate(question[\"criteria\"])}\n        else:\n            raise ValueError(f\"Unknown question type: {question['type']}\")\n        if not 2 <= len(options) <= 20:\n            raise ValueError(\"Provide 2 to 20 criteria per question.\")\n        letters = \"ABCDEFGHIJKLMNOPQRST\"[:len(options)]\n        # Ask for a single option letter, so its logprob represents that option.\n        instructions = question[\"instructions\"]\n        if not isinstance(instructions, str):\n            instructions = json.dumps(instructions)\n        lines = [f\"State:\\n{state}\\n\\nQuestion: {instructions}\\nOptions:\"]\n        for letter, (key, description) in zip(letters, options.items()):\n            line = f\"[{letter}] {key}\"\n            if description is not None:\n                line += f\": {description}\"\n            lines.append(line)\n        prompt = \"\\n\".join(lines) + \"\\n\\nAnswer with the letter of the best option only.\"\n        # OpenAI needs Responses for enough alternatives; llama.cpp needs Chat for logprobs.\n        # top_p=1 avoids pruning alternatives.\n        if is_openai:\n            endpoint = \"/responses\"\n            content = [{\"type\": \"input_text\", \"text\": prompt}]\n            content.extend({\"type\": \"input_image\", \"image_url\": image} for image in images)\n            body = {\n                \"model\": model,\n                \"input\": [{\"role\": \"user\", \"content\": content}],\n                \"reasoning\": {\"effort\": \"none\"},\n                \"max_output_tokens\": 16,\n                \"top_p\": 1,\n                \"top_logprobs\": 20,\n                \"include\": [\"message.output_text.logprobs\"],\n            }\n        else:\n            endpoint = \"/chat/completions\"\n            content = [{\"type\": \"text\", \"text\": prompt}]\n            content.extend({\"type\": \"image_url\", \"image_url\": {\"url\": image}} for image in images)\n            body = {\n                \"model\": model,\n                \"messages\": [{\"role\": \"user\", \"content\": content}],\n                \"max_completion_tokens\": 1,\n                \"temperature\": 0,\n                \"reasoning_effort\": \"none\",\n                \"logprobs\": True,\n                \"top_logprobs\": 1024,\n            }\n        # Send the request and read the first output token's alternatives.\n        request = urllib.request.Request(\n            url.rstrip(\"/\") + endpoint,\n            headers=headers,\n            data=json.dumps(body).encode(),\n        )\n        with urllib.request.urlopen(request) as response:\n            result = json.load(response)\n        if is_openai:\n            message = next(item for item in result[\"output\"] if item[\"type\"] == \"message\")\n            candidates = message[\"content\"][0][\"logprobs\"][0][\"top_logprobs\"]\n        else:\n            candidates = result[\"choices\"][0][\"logprobs\"][\"content\"][0][\"top_logprobs\"]\n        logprobs = {item[\"token\"]: item[\"logprob\"] for item in candidates}\n        # Normalize the returned option scores; missing options initially get zero.\n        missing = [letter for letter in letters if letter not in logprobs or logprobs[letter] <= -9999]\n        if len(missing) == len(letters):\n            raise ValueError(\"API did not return usable scores for any option\")\n        peak = max(logprobs[letter] for letter in letters if letter not in missing)\n        weights = [math.exp(logprobs[letter] - peak) if letter not in missing else 0 for letter in letters]\n        total = sum(weights)\n        # An omitted token cannot outrank the last returned alternative.\n        # Allow zero only when their combined normalized probability is below 1e-6.\n        if missing:\n            cutoff = min(value for value in logprobs.values() if value > -9999)\n            missing_weight = len(missing) * math.exp(cutoff - peak)\n            if missing_weight / (total + missing_weight) >= 1e-6:\n                raise ValueError(f\"API omitted non-negligible option scores for: {', '.join(missing)}\")\n        probabilities = {key: weight / total for key, weight in zip(options, weights)}\n        # Return the winning choice, probability of true, or expected ordinal level.\n        if question[\"type\"] == \"choice\":\n            answers[name] = {\n                \"type\": \"choice\",\n                \"choice\": max(probabilities, key=probabilities.get),\n                \"probabilities\": probabilities,\n            }\n        elif question[\"type\"] == \"noul\":\n            answers[name] = {\"type\": \"noul\", \"noul\": probabilities[\"true\"]}\n        else:\n            answers[name] = {\n                \"type\": \"score\",\n                \"score\": sum(int(key) * probability for key, probability in probabilities.items()),\n                \"legend\": options,\n                \"probabilities\": probabilities,\n            }\n    return {\"answers\": answers}\n# Choose the server and model before opening the camera.\nparser = argparse.ArgumentParser(description=__doc__)\nparser.add_argument(\"url\", nargs=\"?\", default=\"http://localhost:8060/v1\")\nparser.add_argument(\"model\", nargs=\"?\", default=\"gemma-4-12b\")\nargs = parser.parse_args()\n# Point OpenCV's bundled Qt at the installed system fonts.\nos.environ[\"QT_QPA_FONTDIR\"] = \"/usr/share/fonts/truetype/noto\"\n# Open the default Linux webcam with a small capture buffer.\ncamera = cv2.VideoCapture(0, cv2.CAP_V4L2)\nif not camera.isOpened():\n    raise RuntimeError(\"Could not open /dev/video0\")\ncamera.set(cv2.CAP_PROP_BUFFERSIZE, 1)\nprint(f\"Webcam -> {args.model}. Noul: yes %; score: value/max. Ctrl-C or Esc to stop.\", flush=True)\nprint(f\"{'time':<8}\" + \"\".join(f\"{name:>10}\" for name in data[\"questions\"]) + f\"{'fps':>10}\", flush=True)\n# Preview continuously while a background worker scores one frame at a time.\nexecutor = concurrent.futures.ThreadPoolExecutor(max_workers=1)\npending = None\ntry:\n    while True:\n        ok, frame = camera.read()\n        if not ok:\n            raise RuntimeError(\"Could not read a webcam frame\")\n        cv2.imshow(\"Webcam\", frame)\n        if cv2.waitKey(1) == 27 or cv2.getWindowProperty(\"Webcam\", cv2.WND_PROP_VISIBLE) < 1:\n            break\n        # Print a completed result, then submit the latest frame.\n        if pending is not None:\n            if not pending.done():\n                continue\n            result = pending.result()\n            columns = []\n            for name in data[\"questions\"]:\n                answer = result[\"answers\"][name]\n                if answer[\"type\"] == \"noul\":\n                    value = f\"{answer['noul']:.1%}\"\n                elif answer[\"type\"] == \"choice\":\n                    value = answer[\"choice\"]\n                else:\n                    value = f\"{answer['score']:.2f}/{len(data['questions'][name]['criteria']) - 1}\"\n                columns.append(f\"{value:>10}\")\n            columns.append(f\"{1 / (time.perf_counter() - started):>10.2f}\")\n            print(captured + \"\".join(columns), flush=True)\n        # Measure throughput for evaluated frames, including image encoding.\n        started = time.perf_counter()\n        captured = datetime.datetime.now().strftime(\"%H:%M:%S\")\n        ok, jpeg = cv2.imencode(\".jpg\", frame)\n        if not ok:\n            raise RuntimeError(\"Could not encode the webcam frame\")\n        image = \"data:image/jpeg;base64,\" + base64.b64encode(jpeg.tobytes()).decode()\n        data[\"attachments\"] = [image]\n        pending = executor.submit(score, data, args.url, args.model)\nexcept KeyboardInterrupt:\n    print(\"\\nStopped.\")\nfinally:\n    camera.release()\n    cv2.destroyAllWindows()\n    executor.shutdown()\n```\nThe script handles the API differences: llama.cpp uses Chat Completions and OpenAI uses Responses to get it to show alternatives.\n\nI ran Gemma 4 12B QAT through llama.cpp. On Linux with NVIDIA drivers, `curl`, `zstd`, and `uv` installed:\n\n# Model (~7 GB) and multimodal projector (~175 MB).\nmkdir -p ~/models/gemma-4-12b/\ncd ~/models/gemma-4-12b/\ncurl -fL -C - -o gemma-4-12b-it-qat-q4_0.gguf https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-gguf/resolve/main/gemma-4-12b-it-qat-q4_0.gguf\ncurl -fL -C - -o mmproj-gemma-4-12b-it-qat-q4_0.gguf https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-gguf/resolve/main/mmproj-gemma-4-12b-it-qat-q4_0.gguf\n# Standalone llama.cpp binary for RTX 3090 (CUDA architecture 86).\ncurl -fL -o llama.zst https://huggingface.co/buckets/ggml-org/install.sh/resolve/b11160/x86_64/linux/cuda/86/llama-app.zst\nmkdir -p ~/bin/\nzstd -d llama.zst -o ~/bin/llama\nchmod +x ~/bin/llama\n~/bin/llama serve --models-dir ~/models/ --port 8060\n\nSave the Python example as `webcam.py`. In another terminal, from that directory:\n\nuv run webcam.py http://localhost:8060/v1 gemma-4-12b\n# Or use OpenAI, with OPENAI_API_KEY set in your environment.\nuv run webcam.py https://api.openai.com/v1 gpt-6-luna\n","body_html":"<h1 id=\"a-jev-like-wrapper-for-llms-including-vision-models\">A Jev-like wrapper for LLMs, including vision models</h1>\n<p>I was intrigued by <a href=\"https://docs.typesafe.ai/introduction\" rel=\"nofollow ugc noopener\">Jev</a> and the self-hostable projects appearing around it, such as <a href=\"https://huggingface.co/openjev/openjev\" rel=\"nofollow ugc noopener\">OpenJev</a> and <a href=\"https://github.com/TheoLeeCJ/SemIf-OpenJev\" rel=\"nofollow ugc noopener\">SemIf</a>. Reading about them introduced me to a neat trick: reading an LLM&#39;s token probabilities.</p>\n<p>Apparently this is an old trick for some people. See e.g. <a href=\"https://developers.openai.com/cookbook/examples/using_logprobs\" rel=\"nofollow ugc noopener\">OpenAI&#39;s logprobs cookbook</a>. But it was new to me.</p>\n<p>I believe the basic idea is to write a prompt like this:</p>\n<p>State: My order arrived broken and I want a refund.\nQuestion: Which team should handle this?\n[A] billing\n[B] shipping\n[C] returns\nAnswer with the letter of the best option only.</p>\n<p>Then add a few JSON request parameters to a compatible Chat Completions request:</p>\n<pre><code>{\n  &quot;max_completion_tokens&quot;: 1,\n  &quot;logprobs&quot;: true,\n  &quot;top_logprobs&quot;: 20\n}</code></pre>\n<p>The LLM API will return the letter plus the model&#39;s log probabilities for alternative tokens.</p>\n<p>Repeat for each question. Forcing it to generating only one token avoids a lengthy answer and is super quick, though processing the input still costs time. Though for each of the questions a shared state prefix can be KV-cached if the backend supports it.</p>\n<p>The fun part: this works with vision models too. Jev&#39;s <a href=\"https://docs.typesafe.ai/introduction/quickstart\" rel=\"nofollow ugc noopener\">documented request format</a> currently describes only text/JSON state. I added an <code>attachments</code> field for images for my local experiments.</p>\n<p>My example captures webcam frames, sends base64 JPEGs, and prints a table: is a person visible, are we indoors or outdoors, and how bright is the scene? With Gemma 4 12B on my RTX 3090, I get around <strong>1 frames per second</strong>, with three questions per frame. I also ran it against OpenAI gpt-6-luna and got around 0.2 FPS. Presumably because I didn&#39;t make any effort to avoid the cost of a separate connection through their system per question per frame.</p>\n<p>Specialized computer vision models surely are much more efficient, but what I like here is the flexibility: change a condition by describing it in plain text.</p>\n<p>Here&#39;s the standalone Python example (OpenCV is just used for convenient access to the webcam, not for any actual computer vision):</p>\n<pre><code>#!/usr/bin/env -S uv run --script\n# /// script\n# dependencies = [&quot;opencv-python&quot;]\n# ///\n&quot;&quot;&quot;Preview and score webcam frames with llama.cpp or OpenAI.\nuv run webcam.py\nuv run webcam.py https://api.openai.com/v1 gpt-6-luna\nOpenAI reads OPENAI_API_KEY.\n&quot;&quot;&quot;\nimport argparse\nimport base64\nimport concurrent.futures\nimport datetime\nimport json\nimport math\nimport mimetypes\nimport os\nimport pathlib\nimport time\nimport urllib.parse\nimport urllib.request\nimport cv2\n# attachments is our custom addition to the Jev request format.\ndata = json.loads(&quot;&quot;&quot;\n{\n    &quot;state&quot;: &quot;Inspect this webcam frame. Judge only what is visibly present.&quot;,\n    &quot;attachments&quot;: [],\n    &quot;questions&quot;: {\n        &quot;person&quot;: {\n            &quot;type&quot;: &quot;noul&quot;,\n            &quot;instructions&quot;: &quot;Is a person visible?&quot;\n        },\n        &quot;plant&quot;: {\n            &quot;type&quot;: &quot;noul&quot;,\n            &quot;instructions&quot;: &quot;Is a plant visible?&quot;\n        },\n        &quot;setting&quot;: {\n            &quot;type&quot;: &quot;choice&quot;,\n            &quot;instructions&quot;: &quot;Where is the camera?&quot;,\n            &quot;criteria&quot;: {\n                &quot;indoors&quot;: null,\n                &quot;outdoors&quot;: null,\n                &quot;unclear&quot;: null\n            }\n        },\n        &quot;light&quot;: {\n            &quot;type&quot;: &quot;score&quot;,\n            &quot;instructions&quot;: &quot;How bright is the scene?&quot;,\n            &quot;criteria&quot;: [\n                &quot;dark&quot;,\n                &quot;dim&quot;,\n                &quot;bright&quot;\n            ]\n        }\n    }\n}\n&quot;&quot;&quot;)\ndef score(data, url, model):\n    state = data[&quot;state&quot;]\n    if not isinstance(state, str):\n        state = json.dumps(state)\n    # Attachments are our extension to the Jev-style request format:\n    # image file paths or base64 data URLs. Load them once for all questions.\n    images = []\n    for attachment in data.get(&quot;attachments&quot;, []):\n        if attachment.startswith(&quot;data:image/&quot;):\n            images.append(attachment)\n            continue\n        path = pathlib.Path(attachment).expanduser()\n        mime_type, _ = mimetypes.guess_type(path)\n        if mime_type not in {&quot;image/png&quot;, &quot;image/jpeg&quot;, &quot;image/webp&quot;, &quot;image/gif&quot;}:\n            raise ValueError(f&quot;Unsupported image file: {path}&quot;)\n        encoded = base64.b64encode(path.read_bytes()).decode()\n        images.append(f&quot;data:{mime_type};base64,{encoded}&quot;)\n    # Send the API key only to OpenAI.\n    is_openai = urllib.parse.urlsplit(url).hostname == &quot;api.openai.com&quot;\n    headers = {&quot;Content-Type&quot;: &quot;application/json&quot;}\n    if is_openai:\n        headers[&quot;Authorization&quot;] = &quot;Bearer &quot; + os.environ[&quot;OPENAI_API_KEY&quot;]\n    answers = {}\n    for name, question in data[&quot;questions&quot;].items():\n        # Represent choices, booleans, and ordinal levels as lettered options.\n        if question[&quot;type&quot;] == &quot;choice&quot;:\n            options = question[&quot;criteria&quot;]\n        elif question[&quot;type&quot;] == &quot;noul&quot;:\n            options = {&quot;true&quot;: None, &quot;false&quot;: None} | question.get(&quot;criteria&quot;, {})\n        elif question[&quot;type&quot;] == &quot;score&quot;:\n            options = {str(i): description for i, description in enumerate(question[&quot;criteria&quot;])}\n        else:\n            raise ValueError(f&quot;Unknown question type: {question[&#39;type&#39;]}&quot;)\n        if not 2 &lt;= len(options) &lt;= 20:\n            raise ValueError(&quot;Provide 2 to 20 criteria per question.&quot;)\n        letters = &quot;ABCDEFGHIJKLMNOPQRST&quot;[:len(options)]\n        # Ask for a single option letter, so its logprob represents that option.\n        instructions = question[&quot;instructions&quot;]\n        if not isinstance(instructions, str):\n            instructions = json.dumps(instructions)\n        lines = [f&quot;State:\\n{state}\\n\\nQuestion: {instructions}\\nOptions:&quot;]\n        for letter, (key, description) in zip(letters, options.items()):\n            line = f&quot;[{letter}] {key}&quot;\n            if description is not None:\n                line += f&quot;: {description}&quot;\n            lines.append(line)\n        prompt = &quot;\\n&quot;.join(lines) + &quot;\\n\\nAnswer with the letter of the best option only.&quot;\n        # OpenAI needs Responses for enough alternatives; llama.cpp needs Chat for logprobs.\n        # top_p=1 avoids pruning alternatives.\n        if is_openai:\n            endpoint = &quot;/responses&quot;\n            content = [{&quot;type&quot;: &quot;input_text&quot;, &quot;text&quot;: prompt}]\n            content.extend({&quot;type&quot;: &quot;input_image&quot;, &quot;image_url&quot;: image} for image in images)\n            body = {\n                &quot;model&quot;: model,\n                &quot;input&quot;: [{&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: content}],\n                &quot;reasoning&quot;: {&quot;effort&quot;: &quot;none&quot;},\n                &quot;max_output_tokens&quot;: 16,\n                &quot;top_p&quot;: 1,\n                &quot;top_logprobs&quot;: 20,\n                &quot;include&quot;: [&quot;message.output_text.logprobs&quot;],\n            }\n        else:\n            endpoint = &quot;/chat/completions&quot;\n            content = [{&quot;type&quot;: &quot;text&quot;, &quot;text&quot;: prompt}]\n            content.extend({&quot;type&quot;: &quot;image_url&quot;, &quot;image_url&quot;: {&quot;url&quot;: image}} for image in images)\n            body = {\n                &quot;model&quot;: model,\n                &quot;messages&quot;: [{&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: content}],\n                &quot;max_completion_tokens&quot;: 1,\n                &quot;temperature&quot;: 0,\n                &quot;reasoning_effort&quot;: &quot;none&quot;,\n                &quot;logprobs&quot;: True,\n                &quot;top_logprobs&quot;: 1024,\n            }\n        # Send the request and read the first output token&#39;s alternatives.\n        request = urllib.request.Request(\n            url.rstrip(&quot;/&quot;) + endpoint,\n            headers=headers,\n            data=json.dumps(body).encode(),\n        )\n        with urllib.request.urlopen(request) as response:\n            result = json.load(response)\n        if is_openai:\n            message = next(item for item in result[&quot;output&quot;] if item[&quot;type&quot;] == &quot;message&quot;)\n            candidates = message[&quot;content&quot;][0][&quot;logprobs&quot;][0][&quot;top_logprobs&quot;]\n        else:\n            candidates = result[&quot;choices&quot;][0][&quot;logprobs&quot;][&quot;content&quot;][0][&quot;top_logprobs&quot;]\n        logprobs = {item[&quot;token&quot;]: item[&quot;logprob&quot;] for item in candidates}\n        # Normalize the returned option scores; missing options initially get zero.\n        missing = [letter for letter in letters if letter not in logprobs or logprobs[letter] &lt;= -9999]\n        if len(missing) == len(letters):\n            raise ValueError(&quot;API did not return usable scores for any option&quot;)\n        peak = max(logprobs[letter] for letter in letters if letter not in missing)\n        weights = [math.exp(logprobs[letter] - peak) if letter not in missing else 0 for letter in letters]\n        total = sum(weights)\n        # An omitted token cannot outrank the last returned alternative.\n        # Allow zero only when their combined normalized probability is below 1e-6.\n        if missing:\n            cutoff = min(value for value in logprobs.values() if value &gt; -9999)\n            missing_weight = len(missing) * math.exp(cutoff - peak)\n            if missing_weight / (total + missing_weight) &gt;= 1e-6:\n                raise ValueError(f&quot;API omitted non-negligible option scores for: {&#39;, &#39;.join(missing)}&quot;)\n        probabilities = {key: weight / total for key, weight in zip(options, weights)}\n        # Return the winning choice, probability of true, or expected ordinal level.\n        if question[&quot;type&quot;] == &quot;choice&quot;:\n            answers[name] = {\n                &quot;type&quot;: &quot;choice&quot;,\n                &quot;choice&quot;: max(probabilities, key=probabilities.get),\n                &quot;probabilities&quot;: probabilities,\n            }\n        elif question[&quot;type&quot;] == &quot;noul&quot;:\n            answers[name] = {&quot;type&quot;: &quot;noul&quot;, &quot;noul&quot;: probabilities[&quot;true&quot;]}\n        else:\n            answers[name] = {\n                &quot;type&quot;: &quot;score&quot;,\n                &quot;score&quot;: sum(int(key) * probability for key, probability in probabilities.items()),\n                &quot;legend&quot;: options,\n                &quot;probabilities&quot;: probabilities,\n            }\n    return {&quot;answers&quot;: answers}\n# Choose the server and model before opening the camera.\nparser = argparse.ArgumentParser(description=__doc__)\nparser.add_argument(&quot;url&quot;, nargs=&quot;?&quot;, default=&quot;http://localhost:8060/v1&quot;)\nparser.add_argument(&quot;model&quot;, nargs=&quot;?&quot;, default=&quot;gemma-4-12b&quot;)\nargs = parser.parse_args()\n# Point OpenCV&#39;s bundled Qt at the installed system fonts.\nos.environ[&quot;QT_QPA_FONTDIR&quot;] = &quot;/usr/share/fonts/truetype/noto&quot;\n# Open the default Linux webcam with a small capture buffer.\ncamera = cv2.VideoCapture(0, cv2.CAP_V4L2)\nif not camera.isOpened():\n    raise RuntimeError(&quot;Could not open /dev/video0&quot;)\ncamera.set(cv2.CAP_PROP_BUFFERSIZE, 1)\nprint(f&quot;Webcam -&gt; {args.model}. Noul: yes %; score: value/max. Ctrl-C or Esc to stop.&quot;, flush=True)\nprint(f&quot;{&#39;time&#39;:&lt;8}&quot; + &quot;&quot;.join(f&quot;{name:&gt;10}&quot; for name in data[&quot;questions&quot;]) + f&quot;{&#39;fps&#39;:&gt;10}&quot;, flush=True)\n# Preview continuously while a background worker scores one frame at a time.\nexecutor = concurrent.futures.ThreadPoolExecutor(max_workers=1)\npending = None\ntry:\n    while True:\n        ok, frame = camera.read()\n        if not ok:\n            raise RuntimeError(&quot;Could not read a webcam frame&quot;)\n        cv2.imshow(&quot;Webcam&quot;, frame)\n        if cv2.waitKey(1) == 27 or cv2.getWindowProperty(&quot;Webcam&quot;, cv2.WND_PROP_VISIBLE) &lt; 1:\n            break\n        # Print a completed result, then submit the latest frame.\n        if pending is not None:\n            if not pending.done():\n                continue\n            result = pending.result()\n            columns = []\n            for name in data[&quot;questions&quot;]:\n                answer = result[&quot;answers&quot;][name]\n                if answer[&quot;type&quot;] == &quot;noul&quot;:\n                    value = f&quot;{answer[&#39;noul&#39;]:.1%}&quot;\n                elif answer[&quot;type&quot;] == &quot;choice&quot;:\n                    value = answer[&quot;choice&quot;]\n                else:\n                    value = f&quot;{answer[&#39;score&#39;]:.2f}/{len(data[&#39;questions&#39;][name][&#39;criteria&#39;]) - 1}&quot;\n                columns.append(f&quot;{value:&gt;10}&quot;)\n            columns.append(f&quot;{1 / (time.perf_counter() - started):&gt;10.2f}&quot;)\n            print(captured + &quot;&quot;.join(columns), flush=True)\n        # Measure throughput for evaluated frames, including image encoding.\n        started = time.perf_counter()\n        captured = datetime.datetime.now().strftime(&quot;%H:%M:%S&quot;)\n        ok, jpeg = cv2.imencode(&quot;.jpg&quot;, frame)\n        if not ok:\n            raise RuntimeError(&quot;Could not encode the webcam frame&quot;)\n        image = &quot;data:image/jpeg;base64,&quot; + base64.b64encode(jpeg.tobytes()).decode()\n        data[&quot;attachments&quot;] = [image]\n        pending = executor.submit(score, data, args.url, args.model)\nexcept KeyboardInterrupt:\n    print(&quot;\\nStopped.&quot;)\nfinally:\n    camera.release()\n    cv2.destroyAllWindows()\n    executor.shutdown()</code></pre>\n<p>The script handles the API differences: llama.cpp uses Chat Completions and OpenAI uses Responses to get it to show alternatives.</p>\n<p>I ran Gemma 4 12B QAT through llama.cpp. On Linux with NVIDIA drivers, <code>curl</code>, <code>zstd</code>, and <code>uv</code> installed:</p>\n<h1 id=\"model-7-gb-and-multimodal-projector-175-mb\">Model (~7 GB) and multimodal projector (~175 MB).</h1>\n<p>mkdir -p ~/models/gemma-4-12b/\ncd ~/models/gemma-4-12b/\ncurl -fL -C - -o gemma-4-12b-it-qat-q4_0.gguf <a href=\"https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-gguf/resolve/main/gemma-4-12b-it-qat-q4_0.gguf\" rel=\"nofollow ugc noopener\">https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-gguf/resolve/main/gemma-4-12b-it-qat-q4_0.gguf</a>\ncurl -fL -C - -o mmproj-gemma-4-12b-it-qat-q4_0.gguf <a href=\"https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-gguf/resolve/main/mmproj-gemma-4-12b-it-qat-q4_0.gguf\" rel=\"nofollow ugc noopener\">https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-gguf/resolve/main/mmproj-gemma-4-12b-it-qat-q4_0.gguf</a></p>\n<h1 id=\"standalone-llama-cpp-binary-for-rtx-3090-cuda-architecture-86\">Standalone llama.cpp binary for RTX 3090 (CUDA architecture 86).</h1>\n<p>curl -fL -o llama.zst <a href=\"https://huggingface.co/buckets/ggml-org/install.sh/resolve/b11160/x86_64/linux/cuda/86/llama-app.zst\" rel=\"nofollow ugc noopener\">https://huggingface.co/buckets/ggml-org/install.sh/resolve/b11160/x86_64/linux/cuda/86/llama-app.zst</a>\nmkdir -p ~/bin/\nzstd -d llama.zst -o ~/bin/llama\nchmod +x ~/bin/llama\n~/bin/llama serve --models-dir ~/models/ --port 8060</p>\n<p>Save the Python example as <code>webcam.py</code>. In another terminal, from that directory:</p>\n<p>uv run webcam.py <a href=\"http://localhost:8060/v1\" rel=\"nofollow ugc noopener\">http://localhost:8060/v1</a> gemma-4-12b</p>\n<h1 id=\"or-use-openai-with-openai-api-key-set-in-your-environment\">Or use OpenAI, with OPENAI_API_KEY set in your environment.</h1>\n<p>uv run webcam.py <a href=\"https://api.openai.com/v1\" rel=\"nofollow ugc noopener\">https://api.openai.com/v1</a> gpt-6-luna</p>","headings":[{"level":1,"text":"A Jev-like wrapper for LLMs, including vision models","id":"a-jev-like-wrapper-for-llms-including-vision-models"},{"level":1,"text":"Model (~7 GB) and multimodal projector (~175 MB).","id":"model-7-gb-and-multimodal-projector-175-mb"},{"level":1,"text":"Standalone llama.cpp binary for RTX 3090 (CUDA architecture 86).","id":"standalone-llama-cpp-binary-for-rtx-3090-cuda-architecture-86"},{"level":1,"text":"Or use OpenAI, with OPENAI_API_KEY set in your environment.","id":"or-use-openai-with-openai-api-key-set-in-your-environment"}]}}