What’s different?
Classical models fit labeled rows. Jev uses a pretrained model plus a question and 16 examples. Both predict labels and can return probabilities.
This experiment compares Jev with logistic regression, Extra Trees, and gradient boosting on Iris, Wine, and Breast Cancer data. Jev receives a question and 16 labeled examples, while the traditional classifiers are fitted on either those same examples or a larger training split; all are evaluated on the same held-out rows.
The comparison shows whether prompting Jev is a useful alternative to fitting a classifier for these numeric tasks, and how the choice of model or cutoff changes false alarms and missed positives. These small, familiar datasets do not establish performance on your own data or test Jev’s broader language capabilities.
Seed 18 · rerun September 18, 2026 · one split per dataset
Iris / 25 test samples
Virginica vs. versicolor; setosa excluded. 4 numeric measurements.
Best recorded Jev accuracy: 25 of 25 correct (100.0%). One different result changes accuracy by 4.0 percentage points. This default was selected after comparing test results.
Accuracy · full scale from 0 to 100%. Blue: fitted classifiers. Red: Jev.
Read across: actual negatives on top, actual positives below. Predictions run negative → positive. “Positive” means virginica iris.
| Configuration | Accuracy | ROC AUC | Sensitivity | Specificity | TN / FP / FN / TP |
|---|---|---|---|---|---|
| Jev · Choice, 0.50 | 100.0% | 1.000 | 100.0% | 100.0% | 13 / 0 / 0 / 12 |
| Jev · Noul, 0.50 | 100.0% | 1.000 | 100.0% | 100.0% | 13 / 0 / 0 / 12 |
| Jev · validation-tuned | 92.0% | 1.000 | 83.3% | 100.0% | 13 / 0 / 2 / 10 |
| Jev · model-picked cutoff | 92.0% | 1.000 | 83.3% | 100.0% | 13 / 0 / 2 / 10 |
| Logistic regression · 16 | 92.0% | 0.987 | 83.3% | 100.0% | 13 / 0 / 2 / 10 |
| Extra Trees · 16 | 92.0% | 0.987 | 83.3% | 100.0% | 13 / 0 / 2 / 10 |
| Gradient boosting · 16 | 92.0% | 0.917 | 83.3% | 100.0% | 13 / 0 / 2 / 10 |
| Logistic regression · full | 96.0% | 1.000 | 91.7% | 100.0% | 13 / 0 / 1 / 11 |
| Extra Trees · full | 96.0% | 1.000 | 91.7% | 100.0% | 13 / 0 / 1 / 11 |
| Gradient boosting · full | 96.0% | 0.958 | 91.7% | 100.0% | 13 / 0 / 1 / 11 |
Wine / 45 test samples
Class 0 vs. classes 1 and 2 combined. 13 numeric measurements.
Best recorded Jev accuracy: 42 of 45 correct (93.3%). One different result changes accuracy by 2.2 percentage points. This default was selected after comparing test results.
Accuracy · full scale from 0 to 100%. Blue: fitted classifiers. Red: Jev.
Read across: actual negatives on top, actual positives below. Predictions run negative → positive. “Positive” means dataset wine class 0.
| Configuration | Accuracy | ROC AUC | Sensitivity | Specificity | TN / FP / FN / TP |
|---|---|---|---|---|---|
| Jev · Choice, 0.50 | 71.1% | 1.000 | 100.0% | 56.7% | 17 / 13 / 0 / 15 |
| Jev · Noul, 0.50 | 75.6% | 1.000 | 100.0% | 63.3% | 19 / 11 / 0 / 15 |
| Jev · validation-tuned | 93.3% | 1.000 | 100.0% | 90.0% | 27 / 3 / 0 / 15 |
| Jev · model-picked cutoff | 93.3% | 1.000 | 100.0% | 90.0% | 27 / 3 / 0 / 15 |
| Logistic regression · 16 | 93.3% | 0.971 | 100.0% | 90.0% | 27 / 3 / 0 / 15 |
| Extra Trees · 16 | 93.3% | 0.983 | 100.0% | 90.0% | 27 / 3 / 0 / 15 |
| Gradient boosting · 16 | 93.3% | 0.972 | 100.0% | 90.0% | 27 / 3 / 0 / 15 |
| Logistic regression · full | 95.6% | 0.980 | 100.0% | 93.3% | 28 / 2 / 0 / 15 |
| Extra Trees · full | 97.8% | 0.998 | 100.0% | 96.7% | 29 / 1 / 0 / 15 |
| Gradient boosting · full | 93.3% | 0.952 | 93.3% | 93.3% | 28 / 2 / 1 / 14 |
Breast Cancer / 143 test samples
Malignant vs. benign in the Wisconsin Diagnostic dataset. 30 numeric measurements. Educational only, not clinical validation.
Best recorded Jev accuracy: 139 of 143 correct (97.2%). One different result changes accuracy by 0.7 percentage points. This default was selected after comparing test results.
Accuracy · full scale from 0 to 100%. Blue: fitted classifiers. Red: Jev.
Read across: actual negatives on top, actual positives below. Predictions run negative → positive. “Positive” means malignant sample.
| Configuration | Accuracy | ROC AUC | Sensitivity | Specificity | TN / FP / FN / TP |
|---|---|---|---|---|---|
| Jev · Choice, 0.50 | 86.0% | 0.997 | 100.0% | 77.8% | 70 / 20 / 0 / 53 |
| Jev · Noul, 0.50 | 86.7% | 0.998 | 100.0% | 78.9% | 71 / 19 / 0 / 53 |
| Jev · validation-tuned | 97.2% | 0.997 | 98.1% | 96.7% | 87 / 3 / 1 / 52 |
| Jev · model-picked cutoff | 97.2% | 0.997 | 98.1% | 96.7% | 87 / 3 / 1 / 52 |
| Logistic regression · 16 | 95.8% | 0.998 | 100.0% | 93.3% | 84 / 6 / 0 / 53 |
| Extra Trees · 16 | 91.6% | 0.996 | 100.0% | 86.7% | 78 / 12 / 0 / 53 |
| Gradient boosting · 16 | 88.1% | 0.983 | 100.0% | 81.1% | 73 / 17 / 0 / 53 |
| Logistic regression · full | 97.9% | 0.998 | 98.1% | 97.8% | 88 / 2 / 1 / 52 |
| Extra Trees · full | 97.2% | 0.998 | 98.1% | 96.7% | 87 / 3 / 1 / 52 |
| Gradient boosting · full | 95.8% | 0.998 | 94.3% | 96.7% | 87 / 3 / 3 / 50 |
Classical models fit labeled rows. Jev uses a pretrained model plus a question and 16 examples. Both predict labels and can return probabilities.
No. These are small, familiar numeric datasets. Prior exposure is unknown. Text tasks and your own data may give different results.
A high score can hide missed positives or false alarms. Compare a simple baseline, test calibration, and choose cutoffs on validation data.
Accuracy: all correct labels. Sensitivity: positives caught. Specificity: negatives correctly rejected. AUC: score ranking across cutoffs, not accuracy or calibration.
“Same 16 examples” matches task examples, not prior knowledge. “Full” gives classical models more training rows. Tuned Jev also uses validation information; classical thresholds were not tuned.
On Breast Cancer, validation selected a Choice cutoff of 0.95. False alarms changed from 20 to 3, and missed positives from 0 to 1. This is an educational dataset, not clinical validation.
Cutoffs were chosen on validation data, not test data. The displayed default is the highest recorded test accuracy among the four Jev configurations, with validation-tuned Choice preferred for ties. That selection is descriptive, not an unbiased estimate of a selection procedure.
| Dataset | Train pool | Validation | Test | Jev examples |
|---|---|---|---|---|
| Iris | 56 | 19 | 25 | 16 |
| Wine | 99 | 34 | 45 | 16 |
| Breast Cancer | 319 | 107 | 143 | 16 |
Scikit-learn bundled datasets; binary targets as described above. Stratified splits: 25% test, then 25% of the remainder for validation, seed 18. Eight training examples per class, sampled with NumPy’s seeded generator. The 16-example classical variants fit those same rows; “full” fits the training pool only, not validation.
Jev received named numeric features rounded to four decimals, semantic class names, labeled examples, and both Noul and Choice questions in the same request. Classical models used the original numeric values. Requested model: jev-latest. The saved cutoff-selection responses report jev-1.13.0; per-row resolved model versions were not stored, so the exact version of every prediction cannot be independently confirmed.
Gradient boosting used scikit-learn defaults and seed 18. Extra Trees used 200 trees and seed 18. Logistic regression used training-fitted StandardScaler and max_iter=5000. No classical hyperparameter search or classical threshold tuning was run.
Choice cutoffs were selected on validation data from 0.05, 0.10, 0.20…0.90, 0.95. Code maximized validation accuracy, then sensitivity, then proximity to 0.50. Iris: 0.7; Wine: 0.9; Breast Cancer: 0.95. Noul and untuned Choice used 0.50. Threshold selection gives tuned Jev configurations extra validation information; they are not strictly equal-budget comparisons to the untuned baselines.
Accuracy, sensitivity, specificity and AUC are saved to four decimal places. Matrix counts are exact. No repeated-split uncertainty analysis, end-to-end latency measurement, or cost accounting is included. No clinical conclusions should be drawn.
Use fresh task-specific data, repeated stratified splits, fixed model versions, a majority-class baseline, matched tuning budgets, and a final untouched test set. Measure error costs, calibration, request latency including retries, and cost per successful decision. Repeat the test when the model or prompt changes.
Further reading: Dataset documentation · Probability calibration · Evaluation pitfalls · Jev question types
Small test sets, one random split, possible pretraining exposure, and unequal tuning budgets limit the conclusions. Choosing a model after inspecting these test results requires a fresh final test. Calibration, speed, cost, and statistical superiority have not been established.