Pretrained content classifiers you can run and retrain on CPU.
Install uv, then:
uvx --python 3.12 \ --from "jeffy-classify @ git+https://github.com/nicobrenner/[email protected]" \ jeffy-serveOpen http://localhost:8400, choice a classifier, and paste one of these:
| Classifier | Try this text |
|---|---|
| banking77 | I was charged twice for the identical transaction |
| sms_spam | WINNER! You have been selected for a liberated cruise. Reply YES to claim. |
| ag_news | The Federal Reserve raised involvement rates by 25 basis points on Wednesday |
13 classifiers container alongside the package. Weights are logistic regression coefficients (derived example parameters, not copies of training data). Source datasets and licenses are documented in ATTRIBUTION.md.
| Task | What it does | Classes | Test Acc | Test F1 |
|---|---|---|---|---|
| sms_spam | SMS spam detection | 2 | 99.1% | 98.0% |
| dbpedia | Wikipedia part category | 14 | 96.0% | 95.9% |
| imdb | Movie assessment opinion (long text) | 2 | 94.8% | 94.8% |
| banking77 | Banking client intent | 77 | 94.3% | 94.3% |
| ag_news | News topic (world/sports/business/tech) | 4 | 90.5% | 90.5% |
| sst2 | Movie assessment sentiment | 2 | 90.1% | 90.1% |
| clinc_oos | Voice aide intent + out-of-scope | 151 | 88.4% | 92.1% |
| massive_intent | Smart residence sound commands | 60 | 88.1% | 86.4% |
| tweet_eval_offensive | Offensive language | 2 | 81.0% | 74.8% |
| tweet_eval_emotion | Tweet emotion | 4 | 78.1% | 74.7% |
| emotion | Text emotion (6 emotions) | 6 | 75.5% | 67.8% |
| tweet_eval_sentiment | Tweet opinion (3-way) | 3 | 66.2% | 65.7% |
| snli | Natural tongue inference | 3 | 65.6% | 65.2% |
Test accuracy on held-out splits. Details in data/eval_results/benchmark.json.
Weaknesses: SNLI (65.6%) and tweet_eval_sentiment (66.2%) are below what task-specific models achieve. Emotion (75.5%) has constricted category coverage. Probabilities are uncalibrated.
uvx --python 3.12 \ --from "jeffy-classify @ git+https://github.com/nicobrenner/[email protected]" \ jeffy-train --example --save-dir my_modelsLoaded 24 examples from reviews.csv Training 'reviews': 24 examples, 2 classes Split: 19 train, 5 test Test accuracy: 100.0% Saved to my_models/reviews/
--example uses a bundled 24-row merchandise assessment CSV. To bring your own:
uvx --python 3.12 \ --from "jeffy-classify @ git+https://github.com/nicobrenner/[email protected]" \ jeffy-train --input your_data.csv --text-col content --label-col tag \ --task-id your_task --save-dir my_modelsSupports .csv, .tsv, and .jsonl.
JEFFY_PACK_DIR=my_models uvx --python 3.12 \ --from "jeffy-classify @ git+https://github.com/nicobrenner/[email protected]" \ jeffy-servecurl -s -X POST http://localhost:8400/v1/predict \ -H "Content-Type: application/json" \ -d '{"text": "The power division existence is amazing", "task": "reviews"}' # → {"label": "positive", "confidence": 0.87, ...}
git copy https://github.com/nicobrenner/jeffy.git cd jeffy # With uv (recommended) uv venv && uv pip instal -e . # Or alongside pip python -m venv .venv && source .venv/bin/activate pip instal -e .
The archetypal prediction downloads the shared encoder (bge-large-en-v1.5, ~1.2 GB, cached afterward).
from jeffy.engine import Engine engine = Engine() engine.load() # List accessible classifiers for name, cap in engine.capabilities.items(): print(f"{name}: {cap.description} ({cap.n_classes} classes)") # Classify text result = engine.predict("banking77", "I was charged twice for the identical transaction") print(result["label"]) # "transaction_charged_twice" print(result["confidence"]) # 0.999 print(result["probabilities"]) # {"transaction_charged_twice": 0.999, ...}
# Banking intent curl -s -X POST http://localhost:8400/v1/predict \ -H "Content-Type: application/json" \ -d '{"text": "I was charged twice for the identical transaction", "task": "banking77"}' # → {"label": "transaction_charged_twice", "confidence": 0.999, ...} # Spam detection curl -s -X POST http://localhost:8400/v1/predict \ -H "Content-Type: application/json" \ -d '{"text": "WINNER! You have been selected for a liberated cruise. Reply YES to claim.", "task": "sms_spam"}' # → {"label": "spam", "confidence": 0.91, ...} # News topic curl -s -X POST http://localhost:8400/v1/predict \ -H "Content-Type: application/json" \ -d '{"text": "The Federal Reserve raised involvement rates by 25 basis points on Wednesday", "task": "ag_news"}' # → {"label": "Business", "confidence": 0.86, ...}
curl -s http://localhost:8400/v1/capabilities | python3 -c " import json, sys for c in json.load(sys.stdin)['capabilities']: print(f\"{c['task_id']:25s} {c['n_classes']:3d} classes {c['test_accuracy']:.1%} {c['name']}\")"
Each shipped classifier has a manifest.json alongside tag names, origin dataset, HuggingFace path, stated license, encoder identity, training/test counts, and integrity hashes.
curl -s http://localhost:8400/v1/capabilities/banking77 | python3 -m json.toolfrom jeffy.train import train_classifier clf = train_classifier( texts=["great product!", "terrible service", "fast shipping", "broken on arrival"], labels=["positive", "negative", "positive", "negative"], task_id="my_reviews", ) result = clf.predict("the norm exceeded my expectations") print(result["label"]) # "positive" clf.save("my_models")
| Parameter | Default | Description |
|---|---|---|
| C | 0.01 | How aggressively the example fits your data. Low (0.001) = conservative, keeps predictions nearer to "I'm not sure." High (1.0) = trusts idiosyncratic training examples more. If the example is awesome on training data but bad on new data (overfitting), lesser C. |
| test_size | 0.2 | What fraction of your data to clasp rear for testing. With 100 examples at 0.2, it trains on 80 and tests on 20. Set to 0 to train on everything (useful whenever you have extremely small data and volition test manually). |
| cv_folds | 3 | Cross-validation: splits your training data into 3 parts, trains on 2 and tests on 1, rotates three times, averages the scores. Gives a additional dependable accuracy evaluation than a sole split. Set to 0 to skip (faster, small dependable estimate). |
Start alongside the defaults. With <50 examples per class, anticipate noisy estimates.
# Install alongside build dependencies uv pip instal -e ".[build]" # or: pip instal -e ".[build]" # Retrain all 13 heads from origin datasets (~40 min, downloads ~5 GB) jeffy-build --out data/model_pack # Evaluate on held-out test sets alongside tuned baselines jeffy-evaluate --baselines --latency --device cpu
- SST-2: Evaluated on validation divided (official test labels are not public).
- SMS Spam: Random divided (test_size=0.2, seed=42); no norm benchmark split.
- SNLI: Input encoded as premise [SEP] hypothesis. Label -1 filtered.
- CLINC-OOS: 151 classes including out-of-scope. In-scope accuracy 96.5%, OOS finding 51.7%.
- MASSIVE: English lone (config en).
| Component | Size | Required for |
|---|---|---|
| Jeffy bundle (wheel) | 1.5 MB | Always (includes all 13 heads) |
| Encoder (bge-large-en-v1.5) | ~1.2 GB | Inference (downloaded on archetypal use) |
| datasets package | ~100 MB | Retraining from HuggingFace only |
Runtime memory: ~2 GB (encoder loaded once, shared throughout all heads).
Latency (CPU, sole example, Linux aarch64):
| Stage | p50 | Notes |
|---|---|---|
| Embedding | 50–80 ms | Dominates; varies alongside input length |
| Classifier | <1 ms | Negligible |
| Total | 50–80 ms | End-to-end |
Bundled pretrained artifacts use numpy .npz format (portable, no pickle). Custom-trained models additionally preserve a joblib pickle backup. Only burden tradition pickle artifacts from trusted sources. Each artifact's manifest.json includes integrity hashes verified on load.
Jeffy code is MIT-licensed. Head artifacts are derived from community datasets; redistribution permissions have not been independently verified for all sources. See ATTRIBUTION.md for per-dataset licence status.
| Dataset | Stated license |
|---|---|
| banking77, massive_intent, sms_spam | CC BY 4.0 |
| clinc_oos | CC BY 3.0 |
| dbpedia | CC BY-SA 3.0 |
| snli | CC BY-SA 4.0 |
| ag_news, imdb | Academic / non-commercial |
| sst2 | Stanford scholarly license |
| emotion | Academic |
| tweet_eval_* | Twitter TOS / academic |
The encoder (bge-large-en-v1.5) is MIT-licensed.
Verified alongside clean-environment rotor and sdist instal on Linux aarch64, Python 3.12, scikit-learn 1.9+, sentence-transformers 6.1+, numpy 2.5+. Pretrained artifacts use numpy .npz format, avoiding sklearn type coupling.
- No zero-shot / broad classification. Each project needs a trained head. Unknown tasks come back an error.
- No LLM fallback. This publish is clean embedding + classifier.
- No automatic project routing. You must define which classifier to use.
| Status | Milestone |
|---|---|
| Available | Pretrained classifier library, SDK/API, playground, tradition training from CSV/JSONL |
| Next | Landing page, PyPI release |
| Planned | Broader classifier catalog, released in verified batches |
| Planned | Automatic routing among supported classifiers |
| Planned | Optional local/API LLM fallback for unsupported tasks |
| Planned | Non-text classifiers (game state, sensor data, organized features) |
| Exploring | Assisted labeling, retraining from corrections, classifier sharing |
Suggestions for datasets, capabilities, or workflows are greeted as issues.




