How the real TypeSafe Jev model was wired into Jev FCPO, what two live FCPO sessions measured, and the defects they exposed.
Jev FCPO is a personal monitoring and paper-trading workstation for Crude Palm Oil Futures (FCPO) on Bursa Malaysia Derivatives. It is not a broker and it does not place live orders. It is a monitoring platform with an autonomous decision layer, built to do three jobs at once:
STRONG_BULLISH, BULLISH_PULLBACK, CHOPPY_SIDEWAYS).The name is deliberate. The interface and the decision loop are modelled on jev-trade, a crypto perpetuals trading tool. Jev FCPO ports that shape to Malaysian palm oil futures, swapping the crypto tick model for Bursa's contract spec and replacing the generative decision logic with TypeSafe Jev.
A retail FCPO trader working a single chart has to hold four timeframes in their head at once: the 30-minute tide, the 5-minute structure, the 1-minute trigger, and the risk maths that ties them together. Under a live session, that is where discipline breaks. Jev FCPO externalises the whole loop. It does not tell you what to feel; it prints the verdict and logs it, so the trader can benchmark their own calls against an unemotional, repeatable process.
The decision model is not asked to freeform. It scores against a fixed 4-step trend SOP drawn from Elder's Triple Screen, Raschke's EMA pullback, Dow structure, and Wilder ATR sizing:
| Step | Timeframe | Question answered |
|---|---|---|
| 1. Macro tide | 30m | Which direction is allowed? Longs only, shorts only, or flat. |
| 2. Structure | 5m | Has a pullback reached the value zone between EMA 9 and EMA 21 without breaking the swing pivot? |
| 3. Trigger | 1m / 3m | Did the fast bar confirm entry (rejection wick, engulfing, volume expansion)? |
| 4. Risk | all | Is there a deterministic stop and at least 1:2 reward-to-risk? |
Two invariants sit above the steps: never trade against the 30m tide, and never open a new position within ten minutes of a session close. Everything the model can say is bounded by those rules.
The build exists to test a specific claim: that a typed, non-generative decision model can hold a repeatable intraday process better than a human under pressure. That is why the engine logs every verdict with its latency and cost, why the paper book models real friction instead of pretending fills are free, and why this report bothers to measure the model rather than assert that it works. The question is not "can Jev produce a signal". The question is whether its signal is stable, affordable, and disciplined enough to be worth acting on. The rest of this report answers that with data from a live session.
Jev FCPO is not a charting toy. It ingests multi-timeframe FCPO bars and asks a decision model a structured question: what is the best immediate action, how good is the setup, and how exhausted is the move? Everything downstream (the paper book, the stop sizing, the R:R filter) hangs off that answer.
The original code shipped an offline keyword simulator (jev-fcpo-sim-v1) because no model key existed. The simulator scored grade by matching words in a text prompt, which is why early sessions reported Grade A on every single bar. That is not a decision engine; it is a lookup table wearing a lanyard.
The first live session replaced it with the real model.
Jev is reachable two ways, and the code prefers them in this order:
| Priority | Path | Auth | Endpoint |
|---|---|---|---|
| 1 | TypeSafe direct | TYPESAFE_API_KEY | api.typesafe.ai/v1/systemone |
| 2 | Jev via OpenRouter (used here) | OPENROUTER_API_KEY | openrouter.ai/api/alpha/decisions |
| 3 | Offline simulator | none | in-process |
~typesafe/jev-latest is a decisions model, not a chat model. Posting it to /chat/completions returns a 400.
400: "~typesafe/jev-latest is a decisions model and cannot be used with the
chat/completions endpoint. Use the /api/alpha/decisions endpoint instead."
The decisions endpoint accepts the exact { model, state, questions } contract the existing code already built for TypeSafe, so the real model drops in without reshaping the schema. Each question (choice, score, noul) must carry an instructions field or the payload is rejected.
{
"model": "typesafe/jev-1.13-20260917",
"answers": {
"action": { "choice": "ENTER_LONG", "probabilities": { "ENTER_LONG": 1, "ENTER_SHORT": 0, "HOLD": 0 }, "confidence": 1 },
"setup_quality": { "score": 2.84, "probabilities": { "2": 0.15, "3": 0.85 }, "confidence": 0.84 },
"exhaustion_risk": { "noul": 0.22 }
},
"usage": { "input_tokens": 442, "output_tokens": 78, "cost": 0.000018564 },
"provider": "TypeSafe",
"id": "gen-dec-1790003587-YJla9sL0ZkB5NbiVD2cv"
}
Latency sample combines 15 live FCPO bar-close evaluations from the first session with a 12-call controlled benchmark across four synthetic market states (aligned-bullish, aligned-bearish, chop, conflict), 3 repeats each. Full data: docs/jev-bench.json. The 95-call instrumented run in section 4.6 is measured separately and is not folded into this sample.
| Statistic | Value (ms) |
|---|---|
| n | 27 |
| min | 291 |
| p50 (median) | 427 |
| mean | 555.0 |
| p90 | 1069 |
| p95 | 1287 |
| max | 1431 |
| stdev | 301.8 |
| CV (stdev/mean) | 0.54 |
The controlled benchmark repeated each of four market states three times. Scores are continuous and near-stable; the action is not, in the one ambiguous regime.
ENTER_LONG three times with a score spread of just 0.02. The conflict state (bearish 30m tide versus a strong bullish 5m/1m trigger) is not reproducible: two runs returned ENTER_LONG and one returned HOLD on a near-identical prompt. Jev is a sampled model, so a signal that sits on a regime boundary can flip between runs.| Market state | Score min | Score max | Spread | Actions (of 3) |
|---|---|---|---|---|
| aligned-bullish | 2.59 | 2.61 | 0.02 | ENTER_LONG ×3 |
| aligned-bearish | 2.13 | 2.21 | 0.08 | ENTER_SHORT ×3 |
| chop | 0.00 | 0.00 | 0.00 | HOLD ×3 |
| conflict | 1.14 | 1.19 | 0.05 | ENTER_LONG ×2, HOLD ×1 |
The night session moved the engine off the laptop tunnel and onto the host that already runs the other fleet services, behind nginx and TLS. This matters more than it sounds. A Cloudflare quick tunnel is bound to a process on a laptop: close the lid, lose the feed. The engine now runs as a supervised service that survives a reboot, and the webhook endpoint is a stable public URL the TradingView alerts can keep pointing at week after week.
The deployment also forced one security change. The webhook accepted any request that simply omitted the secret field - harmless on localhost, an open door on a public endpoint. Anyone could have posted fabricated bars, and every bar triggers a paid model call and can move the paper book. The secret is now required, not optional. Verified: a request with no secret returns 401, the same as one with a wrong secret.
The second change was measurement. Until this session, Jev's verdicts lived in a 50-entry in-memory ring buffer that emptied on every restart, so the obvious question - when Jev said this, what actually happened? - was unanswerable after the fact. The engine now appends every evaluation to a durable log and stores the answer next to the question:
| Group | Captured per evaluation |
|---|---|
| Identity | decision id, timestamp, bar timestamp, symbol, timeframe, session |
| Market at decision time | OHLC, volume, ATR, RSI, and all four sleeve regimes |
| Jev's answer | action, action confidence, full probability vector, setup score and grade, exhaustion risk |
| Provenance | model version, generation id, provider, and whether the call was real or simulated |
| Performance | latency, input/output tokens, and the real per-call cost returned by the API |
| Gate outcome | whether the gate applied, whether it passed, and the exact failing reason |
| Outcome | executed side and fill price, plus a separate close record linked back to the opening decision |
A 95-evaluation run exercised the full path end to end. The caveat belongs before the numbers: this run was driven by replayed bars generated for the test, not by the live market, though all 95 evaluations were real model calls against a synthesised snapshot. It is an integration test with a real decision model, not market evidence, and it is archived separately from the live log so it cannot contaminate calibration data.
| Metric | Value |
|---|---|
| Evaluations | 95 |
| Real model calls (0 simulator fallbacks) | 95 |
| Actions | HOLD 91, ENTER_SHORT 4 |
| Grades | F 63, C 32 |
| Setup score | 0 (n=63), 1 (n=32) - never reached 2 |
| Action confidence | min 0.39, p50 0.84, mean 0.83, max 1.00 |
| Exhaustion risk | 0.14 to 0.23 |
| Latency | min 293, p50 363, p90 481, p99 781 ms |
| Cost | $0.00339 total, $0.0000357 per call |
| Gate outcome | 4 signals, 0 passed, 4 rejected |
| Trades executed | 0 |
The four entry signals were the interesting part, and none was a close call:
| Signal | Confidence | Setup score | Grade | Exhaustion |
|---|---|---|---|---|
ENTER_SHORT | 0.39 | 0 | F | 0.21 |
ENTER_SHORT | 0.48 | 0 | F | 0.23 |
ENTER_SHORT | 0.49 | 0 | F | 0.22 |
ENTER_SHORT | 0.43 | 0 | F | 0.23 |
NO_DATA for the 1m, 3m and 30m sleeves, so Jev was grading a four-step multi-timeframe setup from the 5-minute sleeve alone. The rulebook's first step is the 30-minute macro tide; with no 30-minute data, a textbook Grade A is unreachable by construction, and Jev declined to invent one. The first session, which had all four sleeves connected, produced a genuine spread including Grade C entries. Confirm the 30m sleeve is streaming before drawing any threshold conclusion from the output.
NO_DATA for most of the session and flipped to BEARISH_PULLBACK at 11:30. A single 30m bar is not a trend.| Metric | Value |
|---|---|
| Cost per call | $0.0000318 |
| Cost per 1,000 calls | $0.0318 |
| Mean input tokens | 757 |
| Mean output tokens | 96 |
| Cost for a 6-hour session, 4 sleeves | $0.0036 |
| Success rate | 27/27 (100%) |
The cost figure is the decisive one: a session costs well under a cent. There is no economic case for the offline simulator whenever a key is present.
ATR risk basis. Stops were sized from whichever bar triggered the call (usually the 1m, ATR ≈ 3.7) when CONTEXT.md requires the 5m structure timeframe. Now reads the 5m sleeve's ATR, falling back to the trigger bar.
Swing pivot ignored. The spec's primary stop basis, the 5m swing pivot, was captured in state but never passed to the paper book. Stops are now placed 1 tick beyond the structural swing when that is wider than the ATR stop, preserving 1:2 R:R.
Fractional grade crash. Jev returns continuous scores (2.84). The code indexed a 4-element array with the raw float, yielding undefined then Grade F. Now rounded before grading. Without this fix the benchmark's 2.59-2.61 scores would all have graded F.
Webhook body validation. Alerts configured with a default TradingView condition send their own text body, not the Pine JSON. The server now logs the raw body so misconfiguration is diagnosable rather than silent.
Optional webhook secret. The secret check only ran when a secret was present, so a request that simply omitted the field bypassed it entirely. Harmless on localhost; on a public endpoint anyone could have posted fabricated bars, and every bar triggers a paid model call and can move the paper book. Now mandatory. Found while preparing the deployment, not by a failure.
Decision log lost on restart. Verdicts were held in a 50-entry in-memory buffer and vanished when the process restarted, making it impossible to ask later what Jev said before a given outcome. Replaced with an append-only durable log that also records the gate result and the eventual close.
Theme tokens were hardcoded. Large parts of the interface were built for a dark background only: status pills, the help modal, sleeve labels, and every profit/loss figure. They did not merely look wrong in light theme, they became unreadable - white text on white. All of it now reads from theme-aware tokens, and every probe point measures at a contrast ratio of at least 5.17:1 in both themes.
Chart signals were invisible. The chart plotted filled trades but not Jev's rejected proposals, so on a session where nothing passed the gate the chart appeared to do nothing at all. Rejected signals now render as distinct markers, separated by timeframe so a 30m alert cannot drop a marker into the 1m chart.
choice/score/noul contract. That contract already existed in the codebase; the integration was a transport swap.test_rich_feed.mjs pushes ~4200-price bars into the same in-memory state as live ~4905 bars and fabricates positions. Validate timeframes on ingest, since anything outside 1/3/5/30 is accepted but invisible to /api/status and the UI.