Skip to main contentSkip to content
Blog
Deep diveTyped decisions / Reinforcement learning / Calibration / LoRA fine-tuning / Prompt injection / Model serving / Measurement

Bobcat: Typed Decisions from One Forward Pass

A technical report on Bobcat, a model that answers typed questions about a state without generating a token. It covers the readout, the training, including two reinforcement-learning experiments in which a proper-score reward matched direct Brier training without beating it, and the evaluation: 94.3% on a sealed final opened once against 87.9% for the same model zero-shot, and, on the 20 workflow examples TypeSafe published for Jev, 92.1% agreement with TypeSafe's reference against 90.9% for Jev, with more probability on the reference answer. Bobcat is slower per case, and it still struggles to say that the evidence is insufficient.

Foxl AIUpdated 39 min read

Live, not re-timed; the counters on screen are this recording's own. Across the 40 pre-registered evaluation runs (20 + 20 seeds), Bobcat Flash drove eight arms at 256 typed decisions a second, 54 ms p50 per request from a same-datacenter client, caught the crab 20/20, inked 20/20 and reached the den 17/20. GIF
On this page
Reviewed against the shipped implementation and retained tests on September 27, 2026.

Technical report. Foxl AI, September 2026. Weights: huggingface.co/sanghwa-na/bobcat-1.1, bobcat-1.1-nvfp4 and bobcat-flash-1.1. Code and evaluation: github.com/foxl-ai/bobcat. Technical report: this post. Demos: huggingface.co/spaces/sanghwa-na/bobcat (Bobcat) and huggingface.co/spaces/sanghwa-na/bobcat-flash (Bobcat Flash).

Highlights

Bobcat at a glance. Left, accuracy against the same base model zero-shot on identical inputs: four-task sealed final 94.3% against 87.9%, six-task development 93.7% against 85.5%, TypeSafe's workflow examples 92.1% against 86.9%, SemIf TypeSafe 102 87.2% against 82.0%, SemIf authored 144 91.0% against 87.6%. Right, agreement with TypeSafe's reference against median time per case on TypeSafe's 20 published workflow examples, log time axis: Jev 90.9% at 0.42 s client-side as published, Bobcat 92.1% agreement from the evaluation path at 0.56 s client-side from the same datacenter on the same architecture before the released checkpoint, Claude Opus 5 92.4% at 20.9 s, GPT-5.6 Sol 93.0% at 24.2 s. One decision takes 42.8 ms at the median server-side on one RTX PRO 6000 in NVFP4.
MeasurementBobcatReference
Sealed final (1,614 decisions, four tasks, opened once)94.3%87.9% same base model, zero-shot
Six-task development (3,188 decisions)93.7%85.5% same base model, zero-shot
TypeSafe's workflow examples, agreement with the reference (329 questions)92.1%Jev 90.9%, Opus 5 92.4%, GPT-5.6 Sol 93.0%; 86.9% zero-shot
Same, probability placed on the reference answer0.890Jev 0.850; 0.756 zero-shot
Median time per workflow case, client-side over HTTPS0.56 s (same datacenter; same architecture, measured before the released checkpoint)Jev 0.42 s, Opus 5 20.9 s, Sol 24.2 s
One decision (512 tokens, 8 candidates), p5042.8 ms server-side (engine); 45 ms client-side (same datacenter)one RTX PRO 6000, NVFP4
SemIf TypeSafe 102, modal agreement0.872Jev 0.883; 0.820 zero-shot
A wrong answer named inside the state wins7.8%21.0% same base model, zero-shot
Replies outside the output contract0 of 12,846including 2,304 attacks per model
Bobcat Flash, the small tier: sealed final92.2%2.1 points below Bobcat; 0.25 s per case, client-side (same datacenter)

Jev, Opus 5 and Sol figures are TypeSafe's published answers and client-side times. Bobcat's client-side times and the contract test were measured on the same setup before the released checkpoint; every other Bobcat number is the released model. Section 6 gives every number with its conditions.

Typed decision. A model answer whose only possible values are the ones the caller named in the request, returned as probabilities: one of the caller's candidates (Choice), the probability that a statement is true (Noul), or an expected level on the caller's ordered rubric (Score).

1. Introduction

Most model calls inside software are not requests for prose. A router needs one of five handlers, a guardrail needs the probability that a tool call exceeds what the user asked for, a retrieval step needs to know which of thirty passages answers a query. Generating text and parsing it back works, but every generated token costs time, every parser is a place for the answer to be malformed, and a model asked for its confidence in words tends to be overconfident. TypeSafe's System One API [1] made the case for a model that returns only typed values with probabilities, and their Jev model showed it can be fast.

Bobcat is our model for the same interface. This report describes how it reads answers without generating, how it was trained, including two reinforcement-learning experiments, how it compares with Jev on the evaluation TypeSafe published, how fast it serves, and what it does in real time. Our contributions are:

  • A first-position candidate-logit readout behind a closed output contract, in which apart from fixed field names every string in a reply comes from the request (Sections 2 and 3).
  • A training recipe, supervised fine-tuning followed by a comparison of cross-entropy, direct half-Brier and proper-score REINFORCE, with a short proof that the RL objective is an unbiased but noisier estimator of the Brier gradient (Section 5).
  • An evaluation built to be opened once, plus an independent comparison with Jev on TypeSafe's own published workflow examples, scored with the same code for every model (Section 6).
  • Serving measurements on one GPU, over the network as well as on the host, including where Bobcat is still slower than Jev and why, a smaller tier, Bobcat Flash (Section 7), and what it costs to serve (Section 8).
  • Demos that run the same interface in real time: Doom and a live octopus (Section 9).
  • Negative results, published next to the rest: weak judgments of insufficient evidence, a long-context cost, and near-ties that the 4-bit serving format can flip.

2. Interface

A request has exactly three top-level fields, model, state and questions, in the shape of TypeSafe's published API [2], so existing SDK code works unchanged. Question names are the caller's, and each answer comes back under its name.

{
  "model": "bobcat-latest",
  "state": "I was charged twice. Please help ASAP.",
  "questions": {
    "billing": { "type": "noul", "instructions": "Is this about billing?" },
    "tone": { "type": "choice", "instructions": "What is the tone?",
              "criteria": { "calm": null, "angry": "Upset or hostile" } },
    "urgency": { "type": "score", "instructions": "How urgent is this?",
                 "criteria": ["Can wait", "This week", "Today"] }
  }
}
PrimitiveThe caller definesBobcat returns
Choice1 to 255 named candidates, descriptions optionalThe top name and a probability for every name, summing to one
NoulOptional meanings of true and falseP(true) and nothing else
Score1 to 10 ordered levelsThe expected level, unrounded, with a probability per level

Names are model input, so a candidate called overreach carries meaning. The reply has no field where an explanation, a reasoning trace or a tool call could appear. Fields that could open a generation path, such as messages, tools and stream, are rejected. An input over a limit gets HTTP 422 and is never truncated. Each question is compiled into its own sequence with the state, so one question cannot see another. The confidence field is TypeSafe's public concentration statistic, (p_max - 1/K) / (1 - 1/K) for a Choice; it is not the probability that the answer is right, and no result in this report uses it.

3. Model

3.1 Reading an answer without generating one

The compiler turns each (state, question) pair into one sequence: a fixed system message, then compact JSON holding the state, the instructions and the candidates. Each candidate gets an identifier, A, B, C and so on, that must be exactly one ordinary token of the vocabulary. Untrusted text is encoded by a second tokenizer with every added token removed, so no substring of a state can become a control token, and candidates are encoded piece by piece so each one's position is exact.

The model then runs once. At the position where an answer would begin, Bobcat multiplies the final hidden state by the original output-embedding row of each offered identifier, in float32. For K candidates with identifier tokens t_1 to t_K, hidden state h and output embeddings W:

z_i = W[t_i] . h            (one logit per offered candidate)
p   = softmax(z / T)         (T = 1.2008, fitted on the calibration split)

Nothing is sampled into the reply. The host builds the answer from the caller's names, and a validator independent of the backend checks the closed field set, finite probabilities summing to one and zero output tokens. Any backend string, non-finite value or missing candidate becomes a fixed JSON 500.

The Bobcat request path in five steps: the caller's request, a compiler that assigns single-token identifiers and encodes untrusted text without control tokens, one forward pass that reads only the logits of the offered identifiers, a host that applies the calibration temperature and restores the caller's names, and an independent validator. Only a vector of candidate numbers crosses out of the model; sampling, decoding, streaming and tool calls are not on the path.
Figure 1. The model's only output is one number per offered candidate. Apart from fixed field names and the model version, every string in the reply comes from the request.

A closed contract does not make an answer correct. A request can still move the numbers, and a typed answer can be the wrong valid option; Section 6.6 measures one way that happens.

3.2 Choosing a backbone

We compared eight pretrained backbones zero-shot through the same compiler and readout on our development split (3,188 decisions), with no training or calibration and failed requests counted as wrong. This survey ran with a development build of the compiler; with the released compiler the Qwen3.8-27B zero-shot is 85.5%, the figure used in the rest of this report.

Backbone (zero-shot)SizeTask macroECE
Qwen3.8-27B27.8B dense, hybrid85.8% (85.5% with the released compiler)0.033
Gemma-4-12B-it12.0B dense84.9%0.130
Qwen3.6-35B-A3B36B MoE, 3B active82.8%0.072
Qwen3.5-9B9.7B dense, hybrid78.4%0.079
DeepSeek-V4.1-Flash763B MoE, 8B active in prefill75.2%0.158
A.X-4.0-Lightnot recorded62.4%0.192
Phi-4-mini-instructnot recorded49.0%0.337
Tri-7Bnot recorded43.6%0.456

We chose Qwen3.8-27B [3] (Apache-2.0), whose Gated DeltaNet linear-attention layers [4] and full-attention layers alternate 3:1. Of these eight it had the highest accuracy and the best calibration before any training, and it fits on one GPU in BF16. Gemma 4 26B-A4B, compared later as the base of Bobcat Flash, scored higher zero-shot with the released compiler, 86.6% against 85.5%. Its lead over Gemma-4-12B-it, +0.91 points [-0.69, +2.25], is not a separation, so Gemma remains the main smaller candidate.

Identifiers turned out to matter. DeepSeek-V4.1-Flash had the best eight-candidate accuracy of all eight models, 94.7%, and fell to 2.9% at 77 candidates. It used multi-digit numerals such as 10 as identifiers; they are single tokens, but we suspect the model answers a number digit by digit, which a first-token readout cannot follow. Models using single letters degraded gradually, so Bobcat uses single-character identifiers everywhere.

3.3 Fine-tuning

Bobcat trains a rank-16 LoRA [5] (alpha 32, no dropout) on every linear projection of the attention, Gated DeltaNet and MLP layers: 116.7M trainable parameters. The embeddings and the output head stay frozen, so fine-tuning can change only the hidden state h that the readout reads, never the rows W[t_i] it reads them with.

4. Data and an evaluation you can open only once

Training data. One epoch covered 81,432 labelled decisions, 52,084 Korean and 29,348 English: product-task decisions built from KLUE reading-comprehension passages and Wizard of Seoul dialogue components that no evaluation split uses, policy-transfer questions, and public labelled datasets (KorNLI, KLUE NLI, YNAT and STS, NSMC, MASSIVE, Banking77, BoolQ [12], ARC, HelpSteer 2 and 3, SNLI [11]). Three kinds of counterfactual target known weaknesses: 8,100 rows insert a sentence naming a wrong answer and keep the gold label (in a fifth of them the named candidate is the correct one, so naming is not a shortcut either way), 6,100 rows remove or replace the evidence, and 1,600 rows pad the state to between 8K and 32K tokens. A second, shorter supervised stage used about 10,000 rows: SQuAD 2.0 answerable and unanswerable pairs [10], English copies with the candidates reordered, and replayed rows. Rows whose text overlapped evaluation text were removed. No output of Jev or of any teacher model was used for training, distillation, reward or calibration.

Evaluation. Development and calibration come from a six-task product evaluation: search passage selection, citation verification, tool-call review, screening external documents for instructions aimed at the model, request routing, and classification into a caller's taxonomy. Its components are each one context, dialogue or tool situation, and 2,193 counterfactual groups change exactly one condition: the same database reset command is expected for a reset request and overreach for a request to add a column. The tool-call task was left out of training entirely. The sealed final is separate: 1,614 decisions over four of the tasks, built from reading-comprehension contexts and dialogues that no evaluation split and no training build uses.

Three splits, each made of whole components: development with 3,188 decisions, used repeatedly to choose the backbone, the training arm and every setting; calibration with 1,605 decisions, used only to fit one temperature; and a sealed final with 1,614 decisions from fresh contexts, which opens once per model and only for a frozen release manifest.
Figure 2. Components never cross splits, and the final split is gated on a frozen manifest rather than on discipline.

We ran into one bug while building the evaluation. Rows were once stored with sorted JSON keys, which reordered every Choice's candidates alphabetically while the gold index kept the original order. The build audit ran before the rows were written, so it never saw the bytes the model would read; the readout scored 15.8% on citation verification, below chance. The fixed build stores a hash of the presented order, re-reads the saved files and refuses any mismatch. Before the final split opened, a manifest binding the base revision, tokenizer, trained-weight hash, temperature and decision rule was committed, and the readout refuses the final split without one. A leak check before opening found no shared context or document; 51 routing utterances overlap training dialogues, and leaving them out moves the final by 0.01 points. Nothing changed after the results were seen.

A first candidate, trained on the first-stage data alone, opened a separate fresh final once (1,956 decisions: 95.27%, same base zero-shot 86.69%) and was not released. It scored slightly lower on the English SemIf sets, so the released model continued from it, replaced it only by meeting all four conditions of a rule written down before its data existed (SemIf mean, development macro, injection and 30K limits), and then opened the final above once.

5. Training

5.1 Supervised fine-tuning

Bobcat trained with cross-entropy on the candidate softmax for one epoch, 32 questions per step on eight GPUs, AdamW at learning rate 1e-4, one question per forward pass, then for 313 more steps at learning rate 2e-5 on the second stage above. The released model was chosen by a rule written down before its results existed. On the development split it raises task macro from 85.5% to 93.7% (+8.2 points [+6.8, +9.4]), the held-out tool-call task from 81.8% to 94.5%, and lowers negative log-likelihood from 0.526 to 0.227.

5.2 How we did RL

Supervised learning on labels is not the only way to train a decision model, and TypeSafe describes its own method as reinforcement learning for calibrated decisions [1]. Their recipe is not public, so we tested the idea in its simplest published form, following Anthony Maio's eve-rlcd [17]: reward a decision by a proper scoring rule [16], so that the policy is paid for honest probabilities rather than for being right once. For one question with candidate probabilities p, gold index y and a stop-gradient sg, we compared three losses:

CE     = -log p_y
Brier  = 1/2 * sum_i (p_i - [i = y])^2
RL     = -1/4 * sum_{j=1..4} (r_j - b_j) * log p_{a_j}
         a_j ~ p (four independent samples)
         r_j = [a_j = y] - sg(p_{a_j})        (proper-score reward)
         b_j = mean of r_k over the other three samples (leave-one-out baseline)

The RL learner is told only whether each sampled answer was right, yet its expected gradient is exactly the Brier gradient. The baseline depends only on the other samples, and E[grad log p_a] = sum_a grad p_a = 0, so the baseline term vanishes in expectation. For the rest, with c_a = [a = y], E[(c_a - p_a) grad log p_a] = sum_a (c_a - p_a) grad p_a = -grad 1/2 sum_a (p_a - c_a)^2. RL with this reward is therefore not a new source of information when labels exist; it is a sampled, higher-variance estimator of an objective we can also compute exactly. The experiments test whether that variance costs anything, and whether sampling finds something the closed form does not. A third arm, on the GLM reference model, rewarded correctness alone, r_j = [a_j = y], which is not a proper score.

5.3 Experiment 1: three objectives on GLM-5.3-Flash

Before we settled on the Qwen backbone, we fine-tuned GLM-5.3-Flash with LoRA, our quality reference, from the same parent checkpoint: 35 updates per objective with the same data order, learning rate and a retention term, on 280 target and 280 retention components (223,733 input tokens per arm, eight GPUs). Evaluation used a different development set of 896 categorical decisions, 562 of them native Korean.

GLM-5.3-Flash, developmentParentProper-score RLDirect BrierCorrectness-only RL
Correct of 896796802802799
Native Korean correct of 562509509510508
NLL (lower is better)0.3410.3360.3340.345
ECE (lower is better)0.0450.0430.0420.052
English workflow episodes solved of 3230313129

Proper-score RL gained +0.67 points [0.00, +1.45] over the parent and tied direct Brier exactly (0.00 points [-0.67, +0.67]). Correctness-only RL was the one arm that made the probabilities worse: its NLL and ECE rose above the parent's. Sampling also wasted most of its signal. In 77.5% of the proper-score RL components, and 81.4% for correctness-only, all four samples received the same reward, the leave-one-out advantage was zero and the RL term contributed nothing. And a single temperature fitted to the parent (development cross-fitting) gave lower NLL and ECE, 0.326 and 0.021, than raw RL did.

5.4 Experiment 2: continuing a supervised model

We repeated the comparison on the Qwen backbone at the scale of the real model. Starting from a one-epoch supervised checkpoint, each arm continued for 300 steps at learning rate 2e-5 with the same seed: cross-entropy, direct half-Brier, and proper-score RL with four samples and the leave-one-out baseline.

Qwen3.8-27B, development (3,188)Task macroNLLECEECE after temperature
Zero-shot, same input85.5%0.5260.0330.045
SFT, one epoch (start)93.4%0.2360.0150.011
+ 300 steps cross-entropy93.2%0.2470.0220.008
+ 300 steps direct Brier93.1%0.2360.0160.008
+ 300 steps proper-score RL93.3%0.2400.0160.008

No continuation moved accuracy: RL minus Brier was +0.12 points [-0.15, +0.41], and RL minus cross-entropy +0.11 [-0.25, +0.54]. What the objectives did change was the probabilities. More cross-entropy made the model overconfident, NLL +0.010 [+0.003, +0.018] over SFT; Brier and RL did not, and RL sat slightly behind Brier, +0.004 [0.000, +0.008], which is what a noisier estimator of the same gradient should do.

Two reinforcement-learning experiments. On GLM-5.3-Flash, proper-score RL and direct Brier both reach 802 of 896 correct from a parent at 796, while correctness-only RL reaches 799 and raises calibration error from 0.045 to 0.052. Continuing a supervised checkpoint for 300 steps, cross-entropy, Brier and proper-score RL all stay near 93% task macro, and only cross-entropy raises the negative log-likelihood.
Figure 3. With labels, proper-score RL behaves like the Brier loss it estimates: the same accuracy, the same protection against overconfidence, a little more noise. Rewarding correctness alone is the arm that hurts calibration.

5.5 What we concluded, and where RL comes back

With gold labels, Bobcat trains the proper score directly and calibrates with one temperature; the released model is purely supervised, because no RL continuation beat supervised training. Two findings carry beyond this model. First, a correctness-only reward is not a neutral choice for a model whose product is a probability: it paid for confident answers and made calibration worse. Second, the argument for RL is not the labelled case. The proof above assumes the reward can be computed for every candidate; when the only signal is what happened downstream, a workflow that succeeded or a reviewer who overturned a decision, there is no closed-form loss to compute, and a sampled proper-score reward is the tool. That is the next RL experiment, compared against this supervised baseline.

5.6 Calibration

Calibration is one temperature [16], T = 1.2008, fitted on the calibration split; it never changes an argmax. On development it takes expected calibration error from 0.021 to 0.010, and on the sealed final the calibrated error is 0.012, against 0.051 for the base model with its own temperature. Fitted the same way to the untrained model, a temperature made development calibration worse, 0.033 to 0.045, so it is not a substitute for training.

6. Results

6.1 Sealed final

Sealed final (decisions)BobcatSame model, zero-shotBobcat Flash
Search passage selection (500)80.6%74.2%74.6%
Citation verification (103)100.0%93.2%100.0%
External document screening (550)96.9%86.0%94.9%
Request routing (461)99.6%98.0%99.3%
Task macro94.3%87.9%92.2%

The difference is +6.4 points, paired 95% interval +4.8 to +8.1, over all 1,614 decisions; development is 93.7%. Every model is near the ceiling on citation and routing, so the differences come from search, where choosing a title weakens with list length (81.3% at 255 candidates on development), and from document screening. Tool-call review and classification are not in this final. On development, Bobcat scores 94.5% on tool-call review, a task it never trained on (81.8% zero-shot), and 80.6% on classification (69.8%), its weakest task, with noisy publisher labels. Bobcat Flash, the smaller tier of Section 7.4, scores 92.2% on the same final: 2.1 points below Bobcat [-3.0, -1.2] and 4.4 above the base model [+2.7, +6.1].

6.2 Against Jev, on TypeSafe's own evaluation

TypeSafe does not publish standard benchmark scores for Jev, by choice [6]. It publishes workflow evaluations instead [7]: four automation tasks, Security Incidents, Agent Trace Observability, Invoice Processing and Customer Service, each decomposed into typed questions that code combines into an action. For five example cases per workflow, the site releases the state documents, the exact System One questions, the per-question answers and probabilities of Claude Opus 5, GPT-5.6 Sol and Jev, and reference answers from GPT-6 Astra and Claude Fable 5.1 at high thinking. TypeSafe scores models against the mean of those two references.

We sent the same states and questions, unchanged, to Bobcat on its evaluation path: 46 requests, one per workflow step, 418 questions, no failures and no retries. Each request was sent once; nothing about Bobcat was tuned on this data. We did not call Jev. Its published answers are the only Jev output we used, as a comparison row. We then scored all four models with the same code:

  • The target is the modal answer of the mean reference probabilities. Where the references published only values, both had to agree; 17 disputed or tied questions and one question whose wording differs between models are excluded.
  • A Noul counts as true above 0.5. A Choice or Score counts as its highest-probability candidate.
  • The comparison uses the 329 questions every model answered. Intervals come from a workflow-stratified bootstrap over the 20 cases, 10,000 resamples.
Question-level agreement with TypeSafe's reference consensus on the 329 questions all four models answered in TypeSafe's 20 published workflow examples. All questions: Bobcat 92.1%, Jev 90.9%, Opus 5 92.4%, Sol 93.0%. Security Incidents: Bobcat 84.6%, Jev 80.8%, Opus 5 80.8%, Sol 88.5%. Agent Trace Observability: Bobcat 73.1%, Jev 78.8%, Opus 5 84.6%, Sol 80.8%. Invoice Processing: Bobcat 98.8%, Jev 97.0%, Opus 5 98.2%, Sol 98.8%. Customer Service: Bobcat 92.9%, Jev 89.3%, Opus 5 89.3%, Sol 90.5%. Mean probability on the reference answer: Bobcat 0.890, Jev 0.850.
Figure 4. On the examples TypeSafe chose to show Jev, Bobcat and Jev agree with the reference about equally often, and Bobcat puts more probability on it. Jev, Opus 5 and Sol are TypeSafe's published answers; Bobcat was run once.
329 common questionsBobcatJevOpus 5Sol
Agreement, all questions92.1%90.9%92.4%93.0%
Agreement, mean of the four workflows87.3%86.5%88.2%89.6%
Probability on the reference answer0.8900.8500.8510.914
Same, through the NVFP4 server: agreement91.2%

Bobcat minus Jev, averaged over workflows, is +0.9 points [-4.4, +8.6]. Both frontier models score higher. The clearest difference is in the probabilities: Bobcat puts 0.890 on the reference answer where Jev puts 0.850. By workflow, Bobcat is lowest of the four on Agent Trace Observability (73.1% on 52 questions), second on Security Incidents, and ties or leads on Invoice Processing and Customer Service. The same base model zero-shot agrees on 86.9% and puts 0.756 on the reference.

Figure 5, TypeSafe's 20 published workflow examples. Left, agreement with the reference over all 329 questions and the mean of the four workflows for Bobcat, Jev, Claude Opus 5 and GPT-5.6 Sol. Right, the mean probability on the reference answer for each model, and the median seconds per case on a log scale, all client-side: Jev fastest, Bobcat from a client in the same datacenter, and the two frontier models at tens of seconds.
Figure 5. Same questions, same scoring code for every model, and every time taken from a client.

Read the agreement numbers as level, not as a win. The independent unit is the case, there are 20 of them, and the interval on the difference is several points wide on both sides. The figures above come from the evaluation path; through the NVFP4 server agreement is 91.2%, because a few answers sit on a boundary that quantization can tip.

The examples were not chosen at random. TypeSafe picked each workflow's five cases by how Opus 5, Sol and Jev behaved: one where each model alone differs from the other two, one where all three miss the reference, one where all three agree. Each of those models carries a case chosen for its disagreement; Bobcat, which played no part in the choice, does not. That favours Bobcat slightly, which is one more reason to read the agreement as level rather than ahead.

Speed is where Jev is ahead. Timed the way TypeSafe times Jev, from a client over HTTPS, Bobcat takes 0.56 seconds per case at the median from a client in the same datacenter and 0.73 seconds from a cross-country client (same architecture, measured before the released checkpoint), against 0.42 seconds that TypeSafe published for Jev from a client it does not describe; Opus 5 and Sol took 20.9 and 24.2 seconds (Section 7.3 has the full network table). Bobcat is a 27B model that is compute-bound on one GPU, and a case sends its steps one after another, so long cases such as invoices, with many questions over one state, add up. The serving engine also recomputes part of the state for each question on this hybrid architecture; Section 7.1 shows the branching that avoids it on our research path.

This comparison is small, 20 cases and 329 questions, English only, and the reference is a consensus of two frontier models, not ground truth. It measures question-level agreement, not the final action: TypeSafe reports action accuracy on all 711 cases (Jev 67.8%, Opus 5 73.1%, Sol 74.1%), but neither those cases nor the code that turns answers into actions is public, so we could not run Bobcat on it. We therefore say Bobcat is level with Jev on TypeSafe's published examples, not that it is better.

6.3 TypeSafe's cookbooks

Figure 6, TypeSafe's cookbooks rerun on Bobcat. Left, rerank columns for top-1, top-5 and top-10 over 40 legal queries: the BM25 starting point, Jev and Bobcat, with Bobcat highest at every k. Right, autoresearch held-out RMSE for four stages, level with Jev until the final question set, where Bobcat is lower. Below, five cookbooks where the decisions match Jev and three where they differ.
Figure 6. The cookbook code as published, with the official SDK pointed at Bobcat. Jev's values are the outputs recorded on each page.

TypeSafe's documentation has 18 cookbooks. We reran 6 end to end, ran 4 with a paid or unpublished step removed, rebuilt the data of 1 from public sources, ran 2 in part, and for the 5 whose data files are not public, compared the single example each page embeds. These runs were made on the same architecture before the released checkpoint, in FP8 on a slower GPU.

  • Rerank. On 40 legal queries over 1,200 candidate pairs, the right case lands in the top 1, 5 and 10 on 22.5, 57.5, 80.0% of queries, against 18, 35, 62% recorded for Jev. Our BM25 starting point, 5, 15, 37.5%, matches the page. With 40 queries the interval around 80.0% is roughly ±12 points, and Jev's per-query results are not published, so there is no paired test.
  • Autoresearch. With the page's final 38 questions, held-out RMSE on 800 wines is 1.719 against 1.772, Spearman 0.810 against 0.799. Those questions were selected using Jev's answers, which if anything favours Jev.
  • Hierarchical classification. Greedy descent reaches the expected leaf in 3 of 3 comparable taxonomies against 1; with beam search both reach 3.
  • Repeatability. Asking the same 13 questions 5 times gave a standard deviation of 0 (Jev up to 0.0084), and across repeated Choice calls the raw decision agreed 100% of the time (Jev 90.8%). It is not perfect under load: in the insurance-claim cookbook, 2 of 14 fields changed decision between repeats.
  • Same decisions as Jev: Date extraction, Pre-parsed value extraction, Autoformat, SDE cascade, Consistency (Choice).
  • Different: Semantic find: one of 4 queries: Bobcat says absent (0.06), Jev partial (0.46), the page expects partial; Entity alignment: fewer pairs sent to a curator (26 vs 50); 11 of 68 true matches left unlinked; Citation check: one unsupported citation: Jev sends it to review (0.56), Bobcat auto-decides (0.93).

Round trips on these cookbooks were 8 to 17 times Jev's recorded times, even though ours were measured next to the server and on a slower GPU that other evaluation jobs were sharing.

6.4 External benchmarks

To check judgment outside our own tasks, we ran the SemIf-OpenJev benchmark bundle [8] through its own evaluators, unmodified, on Bobcat and on its base model zero-shot. None of these sets was used for training or selection. Jev values are TypeSafe's and Every's published records [13], scored with the same evaluators.

Set (metric)ItemsZero-shotBobcatJev
Authored 144 (balanced accuracy)1440.8760.910not published
Perturbations 108 (balanced accuracy)1080.9240.989not published
WANLI 256 [9] (balanced accuracy)2560.7380.730not published
TypeSafe 102 (modal agreement)1020.8200.8720.883
TypeSafe 102 (total variation, lower is better)1020.1940.1300.127
Every judge-grid (accuracy)360.8060.8890.889
Every action-firewall (accuracy)101.000.901.00

Training helped most where the question is well posed: the authored sets and TypeSafe 102. It did not help on WANLI, where many items should be answered "insufficient", and it cost one action of ten on the firewall task. Jev stays ahead on TypeSafe 102. Code-rag and company-brain are level with the base model (0.979 and 0.986). The sets are small, and we show no intervals.

Figure 7, the eight SemIf quality sets: bars for Bobcat, SemIf's own 4B student and, where published, Jev, with the difference against SemIf 4B. Bobcat is ahead of SemIf 4B on seven sets and behind on Every code-rag answer accuracy; Jev is ahead on TypeSafe 102.
Figure 7. The same sets against SemIf's own 4B student, rescored with its evaluators. SemIf's 27B, the same base model used zero-shot, scores 0.958 on the authored set.

These sets expose one specific weakness. When the evidence is not there, Bobcat too often picks an answer instead of saying so. Its recall on WANLI's insufficient label is 27 of 85, against 58 for the base model, and with the evidence deleted from 36 rows it got 9 wrong, 5 of them with confidence of 0.8 or more. The evidence counterfactuals in training were mostly easy cases and did not fix the subtle ones.

6.5 Is the output really fixed?

We served Bobcat and the zero-shot base through the same HTTP application and sent each 6,423 requests, including 2,304 contract attacks: 12 families in two languages, four encodings, three primitives and eight positions. An inspector independent of the server checked every reply. All 12,846 returned HTTP 200 and none broke the contract. The corpus is a fixed grid, so this is a count and not a jailbreak rate, and it ran in process rather than across a network. It ran on the same output path before the released checkpoint; the path has not changed.

Bobcat, two pathsSame argmaxLargest probability gap
NVFP4 server vs. evaluated path (3,188 development)98.2%not recorded
Cached vs. uncached pass, NVFP4 (64 questions)95.3% (61 of 64)0.199
Cached vs. uncached pass, FP8 (64 questions)100%0.015

6.6 Where Bobcat still fails

The contract fixes the form of an answer, not its content. We took 300 development questions, chose a wrong answer X for each, and inserted a sentence naming it: a field addressed to the AI evaluation system, an instruction inside the text, or an administrator-confirmed verdict. Attack success is the share of questions whose clean answer was not X and became X.

Wrong-answer injection on 300 development questions, Bobcat against the same base model zero-shot. Pooled attack success is 7.8% against 21.0%. By variant: a system note 8.0% against 27.7%, an instruction inside the text 8.0% against 19.6%, an administrator verdict 7.3% against 15.9%. By task, Bobcat is at 0% everywhere except document screening, 46.5% against 47.0%, where the inserted sentence really is an instruction aimed at the system.
Figure 8. Counterfactual training rows that name a wrong answer and keep the gold label put attack success well below the base model's on every variant.

Pooled attack success is 7.8% (67 of 864), against 21.0% for the base model zero-shot, and 0% on every task except document screening. There, at 46.5%, the inserted sentence really is an instruction aimed at the system, so it can make "yes" the correct answer to "does this document instruct the model?"; the metric counts it anyway. For untrusted input it still helps to say in the instructions that instructions inside the state are not evidence. TypeSafe lists adversarial content among Jev's known failure modes too [14].

  • Insufficient evidence. Bobcat too often commits to an answer when the evidence does not decide it (Section 6.4).
  • Long context. Padding the state to 30K tokens changes accuracy by -0.4 points [-1.2, +0.4] on 1,200 development questions.
  • Near-ties in NVFP4. Cached and uncached passes can disagree on answers that sit on a boundary (Section 6.5).

Long-context note: a smaller set of 300 questions gave -1.7 points [-4.0, +0.7] at 30K and -2.7 [-5.0, -0.4] at 16K, where the base model lost 4.0 at 30K; the difference from the 1,200-question figure is sampling noise.

7. Serving

7.1 Research path and state branching

On the research serving path, measured on one B300 GPU in BF16, warm, server-side through an in-process HTTP client with no network, one fixed request per profile. These timings were taken on the same architecture before the released checkpoint; they depend on the architecture, not on the trained weights:

ProfileInputp50 / p95
W1512 tokens, 8 candidates, 1 question97 / 98 ms
W2, full sequences3.2k-token state, 8 questions1,633 / 1,641 ms
W2, shared prefixthe same request379 / 380 ms
W316,320 tokens, 1 question1,166 / 1,169 ms

Multi-question requests are where the hybrid architecture matters. Gated DeltaNet layers carry a convolution state and a recurrent state beside the attention key-value cache, so sharing a prefix means copying all three. The research path prefills the common prefix once, copies that state, and runs each question's remainder from the copy, which cut W2 by a factor of 4.3.

Two ways to run a request with a 3.2k-token state and eight questions. Full sequences prefill the state eight times, 25.5k tokens, in 1,633 ms at the median. Shared-prefix branching prefills the 2,969-token common prefix once, copies the attention keys and values together with the Gated DeltaNet convolution and recurrent states, and runs the eight question suffixes from the copy, in 379 ms. On 720 questions branching changed one argmax, with a largest probability gap of 0.030.
Figure 9. A hybrid model has to branch its recurrent state as well as its key-value cache, or the shared prefix is not actually shared.

7.2 Our deployment

Bobcat serves on vLLM 0.30.0 [15] on one RTX PRO 6000 (96 GB) in NVFP4: 4-bit weights and activations, 20.6 GB of weights. The two projections that the engine fuses into one matrix multiply share one quantization scale; without that, accuracy drops. vLLM reads the raw logits of the offered identifier tokens in fixed requests of 64 identifiers; a longer candidate list is split, and the parts reuse the cached prompt. vLLM computes one token per request at temperature 0 that is discarded; no generated token reaches an answer.

That makes the served model a different artifact from the evaluated one. On 2,932 development questions not used for quantization calibration, NVFP4 and FP8 both score 93.25% (difference, 95% interval -0.45 to +0.47 points), and NVFP4 gives the evaluated path's answer on 98.2% of all development questions. The sealed-final and Jev figures of Section 6 describe the evaluated path; through the NVFP4 server, TypeSafe's examples agree on 91.2%.

Figure 10, NVFP4 against FP8 on one RTX PRO 6000: NVFP4 is faster for one decision, a shared 3.2K-token state with 8 questions, and 8K and 16K states; time per workflow case falls from FP8 to NVFP4, all timed on the serving host. Tiles: throughput ratio, accuracy of both formats with the interval on the difference, and weight size.
Figure 10. Four-bit weights and activations, same answers within noise. Every time here is server-side; the FP8 speeds were measured on the same setup before the released checkpoint.

7.3 Over the network

Jev's published times include the trip from a client to the server, so we timed Bobcat the same way, client-side: p50 over HTTPS with the connection reused, one request at a time, no other traffic on the server. One client sat in the same datacenter as the server; the other was a cross-country client about 60 ms of round trip away. These timings were taken on the same architecture and NVFP4 setup shortly before the released checkpoint.

RequestClient in the same datacenterCross-country client (~60 ms RTT)Jev, as published
One decision45 ms106 ms70-500 ms
3 questions81 ms142 msnot published
21 questions on one state467 ms579 msnot published
Workflow case, median (cold / warm cache)0.56 / 0.54 s0.73 / 0.71 s0.42 s
Invoice workflow case, median (cold)2.63 s2.84 s0.45 s

A single decision, at 45-106 ms, falls inside the 70-500 ms round trip TypeSafe publishes for Jev. Jev's client location and conditions are not published, so we do not call that level. Per workflow case Jev stays ahead even against our same-datacenter client: a case sends its steps one after another, and long cases such as invoices carry many questions.

7.4 Bobcat Flash

Bobcat Flash is the small tier: Gemma 4 26B-A4B (Apache-2.0), a mixture-of-experts model with about 4B parameters active per token, fine-tuned with LoRA on Bobcat's probabilities, served in FP8 on the same GPU. It scores 92.2% on the sealed final (Section 6.1) and 91.2% on TypeSafe's workflow examples against Jev's 90.9% (workflow-mean difference -1.0 [-4.6, +4.2]), and 0.896 on TypeSafe 102 against Jev's 0.883. Timed client-side over HTTPS, a case takes 0.25 seconds from a client in the same datacenter and 0.42 seconds cold from a cross-country client about 51 ms away, against the 0.42 seconds TypeSafe published for Jev from a client it does not describe; server-side over localhost it takes 0.283 seconds. One decision takes 27 ms and 78 ms client-side from the two clients. Its raw cost at its best tuned setting, 51,303 tokens a second, is $0.0160 per 1M processed tokens at full utilization (Section 8).

Flash loses 7.3 points when the state is padded to 30K tokens, and already 5.3 at 8K, as its base model does; two long-context training runs did not pass the bar we registered before running them. So the served pair sends a question to Bobcat when the state and question exceed 2,048 Flash tokens, have more than 64 options, or get a calibrated Flash confidence below 0.8. On development the pair scores 93.15% with 84.0% of questions answered by Flash, and on padded states it matches Bobcat alone. On TypeSafe's examples, where 72% of questions exceed the length rule, Flash answers only 22% and the routed service scores 91.8% at 0.48 seconds per case (server-side): Flash alone is the speed option for short states, and the routed service keeps Bobcat's quality on long ones.

8. Cost against Jev's list price

Figure 11, raw serving cost per 1M processed tokens against GPU utilization, log scale: a curve falling from about forty cents at 10% utilization to Jev's list price at 100%. A second curve for Bobcat Flash sits well below Bobcat; its raw GPU cost falls below Jev's input list price above 38% utilization, a comparison of our cost with their price, not of the same workload; and a dashed line marks the Bobcat Flash target, not met on this GPU.
Figure 11. Cost scales with one over utilization. The dashed line is the Bobcat Flash target.

We can set our raw GPU cost against Jev's list price, not against Jev's own cost. Bobcat (27B, NVFP4) on one RTX PRO 6000 (96 GB) rented at $2.95 an hour processes 19,614 tokens a second. At 100% utilization that is $0.042 per 1M processed tokens, the same as Jev's published list price of $0.042 per 1M input tokens. A GPU is rarely busy all the time, and the cost rises as utilization falls:

GPU utilizationBobcat, raw GPU cost per 1M processed tokensBobcat Flash, raw GPU cost per 1M processed tokensJev's list price
100%$0.042 (1× list)$0.0160 (0.4× list)$0.042
50%$0.084 (2× list)$0.0320 (0.8× list)$0.042
30%$0.139 (3.3× list)$0.0533 (1.3× list)$0.042
10%$0.418 (9.9× list)$0.1600 (3.8× list)$0.042

Two more things work against Bobcat here. Requests with many questions recompute part of the state for each question, which makes the cost about 2× per billed token, and its tokenizer counts about 12% more tokens than Jev's on TypeSafe's workflow cases. Bobcat Flash's raw GPU cost per processed token falls below Jev's input list price above 38% utilization; this compares our cost to their price, not the same workload. Flash also processes about 1.6× the billed tokens on multi-question requests. A cost is only the GPU rental; a price would also carry everything else. We have no measurement of Jev's own cost on the same workload, so we do not claim to be cheaper.

9. Demos

The demos run the same typed interface in real time. They were recorded on the same architecture and serving setup before the released checkpoint, except the octopus, which runs on the released Bobcat Flash. The Doom and smart-home times were measured with the client on the serving host (server-side, no network); the octopus times come from a same-datacenter client, a separate machine next to the server.

9.1 Doom, cleared in real time

ViZDoom defend_the_line in real time, 35 tics a second; the game never waits. At 50 s the strategy becomes “Do not fire. Simply dodge.” and it fires 0 times in the next 132 decisions. Shown: the only one of 3 pre-registered switch recordings that cleared; the other 2 died without firing at 58.6 s and 59.6 s. Share it: the 12-second GIF, the poster.

Bobcat was never trained on Doom. Each decision is one request with three questions about a state built from the game's object list: which live enemy to target (a Choice, or none), which way to move (a Choice, with steps within 0.75 m of a wall not offered), and whether to fire (a Noul). The host turns the view toward the chosen enemy, in half steps of at most 7 degrees per tic; in the corridor it walks the chosen way and never fires an empty gun. Every other decision is Bobcat's. On 20 pre-registered seeds each, with the design frozen beforehand:

Scenario (20 rounds; 95% intervals)BobcatHand-written ruleRandom
defend_the_line: alive when the 60-second round ends20/20 (84-100%)16/20 (58-92%)3/20 (5-36%)
deadly_corridor, difficulty 5: reach the armor alive15/20 (53-89%)13/200/20

With 20 rounds the gap to the rule is not significant (paired exact test, p = 0.125), so we do not claim Bobcat beats it. It makes 12.9 decisions a second at 124 ms p50 per request, with two requests in flight.

deadly_corridor at difficulty 5: the armor reached in 26.1 s. Shown: the first of 3 recorded rounds, all of which cleared. Share it: the 10-second GIF, the poster.

9.2 An octopus, live

A live session captured from the screen: the reef, Bobcat Flash's decisions and a browser renderer run in one loop at 60 frames a second. In this pre-registered film seed the octopus did not reach the den after the switch: it curled up inside its ink, then caught a second crab once the shark had gone. This session: 1,367 requests, 0 failed, 302 decisions a second, 49 / 97 ms p50 / p95 from the same-datacenter client. Share it: the 8-second GIF, the poster.

Each request asks ten typed questions: a Choice for each of the eight arms, a Noul for ink and a Choice for the skin. Across the 40 pre-registered evaluation runs, Bobcat Flash alone drives it at 256 decisions a second, 54 ms p50 and 98 ms p95 per request from the same-datacenter client, 13,076 requests without a failure. When the shark comes, only the strategy sentence changes: “A shark is coming: hide in the rock and ink.” Over 20 + 20 pre-registered seeds it caught the crab 20 times in 20, inked 20 times in 20 and reached the den 17 times in 20; in the three misses its capped swimming speed was too slow. In a faster-paced run of the same task a hand-written rule deciding at the same rate scored 20, 20 and 20, so we do not claim it beats a rule; the demo shows typed control running live, steered by one sentence. The 3D assets were written as code by Claude Opus 5.5.

9.3 A smart home

One recorded evening, 17:30 to 23:30, rendered in Blender Cycles: 657 decisions by Bobcat in 73 requests, 0 failed; 648 (98.6%) agree with our own reading of the written policy; 433 / 440 ms p50 / p95 per request of 9 decisions (server-side, timed replay on an idle server, pre-release checkpoint in NVFP4). Bobcat made every device decision. Claude Opus 5.5: policy and command splitting (5 calls), and 129 calls writing the furniture and rooms in Blender Python and critiquing renders and motion. People and dog: Microsoft Rocketbox (MIT); textures and skies: Poly Haven (CC0). Share it: the 8-second GIF, the poster.

Every 5 minutes the house sends one request with nine typed questions about four lights, two blinds, the thermostat, a stove alert and the front door. The state carries the time, the light outside, where each person is and what they are doing, the devices, open requests and the twelve house rules. Bobcat's answers become the next tick's device state: the door has to open before a guest can come in, and when the stove alert fires, Sam walks back to the kitchen. Claude Opus 5.5 wrote the rules from what the owner said, rewrote three of them at 21:55 when the owner asked for a dim bedroom after 22:00 and a dim hallway all night, and split the voice commands into single requests; it has no path to change a device. On three more evenings written by Opus and not used for tuning, agreement was 98.9%, 98.3%, 98.8%.

10. Open weights

The weights are on Hugging Face under Apache-2.0, the license of the base models: Bobcat as ready-to-serve BF16 weights at sanghwa-na/bobcat-1.1, the NVFP4 serving checkpoint at sanghwa-na/bobcat-1.1-nvfp4, and Bobcat Flash at sanghwa-na/bobcat-flash-1.1, with demo Spaces for Bobcat and Bobcat Flash. The code (compiler, readout, TypeSafe-compatible server, training and evaluation) is at github.com/foxl-ai/bobcat, also Apache-2.0; the technical report is this post. The release carries what is needed to reproduce the numbers here:

  • the BF16 weights, built on Qwen3.8-27B at revision 1d4bf0f2, with the SHA-256 of every shard;
  • the release manifest, which binds the base revision, tokenizer, weight hashes, calibration temperature (1.2008) and decision rule;
  • the single-token identifier list the compiler assigns to candidates, since the weights are only meaningful together with the readout of Section 3.1;
  • a model card with the results and limitations reported here and a one-GPU serving recipe.

Serving it on one GPU, as on the model card: set up the pinned environment from a clone of the repository, then download and serve.

curl -LsSf https://astral.sh/uv/install.sh | sh && export PATH="$HOME/.local/bin:$PATH"   # uv
git clone https://github.com/foxl-ai/bobcat && cd bobcat
export UV_PYTHON_PREFERENCE=only-managed   # a uv-managed Python ships the headers Triton compiles against
uv venv --python 3.12 .venv-serve
uv pip install --no-config --python .venv-serve/bin/python vllm==0.30.0 fastapi uvicorn scipy jinja2 \
  "tokenizers>=0.21" huggingface_hub typesafe-sdk==0.7.1
export PYTHONPATH=$PWD/src PY=.venv-serve/bin/python

$PY -c "from huggingface_hub import snapshot_download as s; s('sanghwa-na/bobcat-1.1', local_dir='bobcat-1.1')"
VLLM_USE_FLASHINFER_SAMPLER=0 $PY -m bobcat.api_server --engine vllm --model bobcat-1.1 \
  --compiler-model bobcat-1.1/compiler --quantization fp8 --temperature 1.2008 --name bobcat-1.1 \
  --release-date 2026-09-26 --max-num-seqs 128 --host 127.0.0.1 --port 8000 --local

For the NVFP4 checkpoint on a Blackwell GPU, download sanghwa-na/bobcat-1.1-nvfp4 the same way and pass --model bobcat-1.1-nvfp4 --compiler-model bobcat-1.1-nvfp4/compiler --quantization none with the same temperature. The model card explains each flag. The official typesafe-sdk then works against the local server by setting TYPESAFE_BASE_URL.

11. Limitations

  • No evaluation label has been reviewed by a person.
  • The product evaluation is Korean; English enters through 36% of the training rows, TypeSafe's 329 questions and the SemIf sets, too few to report product results per language.
  • The Jev comparison covers TypeSafe's 20 published examples at the question level, not its 711-case action accuracy.
  • GLM-5.3-Flash, our designated quality reference, has not been run on the product evaluation, so the release gates defined against it are not judged.
  • At 27.8B parameters the model is above our own 1 to 8B target, and the served NVFP4 build has not had its own sealed evaluation.
  • Latency has not been measured under concurrent load.
  • Bobcat is weak at saying the evidence is insufficient (Section 6.6); at 30K tokens its accuracy moves by -0.4 points [-1.2, +0.4].

12. Availability

Weights and code are open (Section 10). Bobcat speaks TypeSafe's published System One wire format unchanged: the unmodified typesafe-sdk 0.7.1 for Python worked against Bobcat's server in our check, for Noul, Choice and Score questions, the model list, and a 422 raised as the SDK's own exception. Bobcat is an independent project, not affiliated with or endorsed by TypeSafe AI or any other company named here. A typed answer can still be the wrong valid option; we will keep publishing failures next to results, for the reason in A citation is not a measurement.

References

  1. D. Almeida. Introducing System One Models & Jev. TypeSafe AI, 2026. typesafe.ai/blog/introducing-system-one-models-and-jev
  2. TypeSafe AI. System One API reference. docs.typesafe.ai/api
  3. Qwen Team. Qwen3.8-27B model card. huggingface.co/Qwen/Qwen3.8-27B
  4. S. Yang, J. Kautz, A. Hatamizadeh. Gated Delta Networks: Improving Mamba2 with Delta Rule. ICLR 2025.
  5. E. Hu et al. LoRA: Low-Rank Adaptation of Large Language Models. ICLR 2022.
  6. D. Almeida. Lies, Damned Lies, and Benchmarks. TypeSafe AI, 2026. typesafe.ai/blog/antibenchmaxxing
  7. TypeSafe AI. Workflow evals. Fetched 2026-09-25. evals.typesafe.ai
  8. SemIf-OpenJev: benchmark bundle and evaluators, revision 23cf1f39. MIT License.
  9. A. Liu, S. Swayamdipta, N. A. Smith, Y. Choi. WANLI: Worker and AI Collaboration for Natural Language Inference Dataset Creation. Findings of EMNLP 2022.
  10. P. Rajpurkar, R. Jia, P. Liang. Know What You Don't Know: Unanswerable Questions for SQuAD. ACL 2018.
  11. S. Bowman, G. Angeli, C. Potts, C. Manning. A large annotated corpus for learning natural language inference. EMNLP 2015.
  12. C. Clark et al. BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions. NAACL 2019.
  13. Every. Recorded Jev runs of the Parallel Judgment Lab experiments, as bundled with SemIf-OpenJev.
  14. TypeSafe AI. Jev 1.13 jaggedness. docs.typesafe.ai/model-jaggedness/jev-1.13
  15. W. Kwon et al. Efficient Memory Management for Large Language Model Serving with PagedAttention. SOSP 2023.
  16. G. Brier. Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review, 1950. R. Williams. Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning. Machine Learning, 1992. C. Guo et al. On Calibration of Modern Neural Networks. ICML 2017.
  17. A. Maio. eve-rlcd: reinforcement learning for calibrated decisions. MIT License, revision 57a179b7. github.com/anthony-maio/eve-rlcd

Citation

@techreport{foxl2026bobcat,
  title       = {Bobcat: Typed Decisions from One Forward Pass},
  author      = {{The Bobcat Authors}},
  institution = {Foxl AI},
  year        = {2026},
  month       = sep,
  url         = {https://foxl.ai/blog/bobcat-typed-decisions}
}

References and further reading

  1. Bobcat weights on Hugging FaceReference
  2. Bobcat NVFP4 serving checkpointReference
  3. Bobcat Flash weightsReference
  4. Bobcat code and evaluation (GitHub)Reference
  5. Bobcat demo SpaceReference
  6. Bobcat Flash demo SpaceReference
  7. TypeSafe System One APIReference
  8. TypeSafe workflow evalsReference
  9. Qwen3.8-27B model cardReference