Case Study · Technical Report

Using TypeSafe Jev as the Decision Engine for Jev FCPO

Date: 2026-09-21 Sessions: BMD T-Morning + T+1 Night Model: typesafe/jev-1.13-20260917

How the real TypeSafe Jev model was wired into Jev FCPO, what two live FCPO sessions measured, and the defects they exposed.

1. Background: what Jev FCPO is, and why we built it

Jev FCPO is a personal monitoring and paper-trading workstation for Crude Palm Oil Futures (FCPO) on Bursa Malaysia Derivatives. It is not a broker and it does not place live orders. It is a monitoring platform with an autonomous decision layer, built to do three jobs at once:

  1. Watch the market across four timeframes. It receives live FCPO bars (1m, 3m, 5m, 30m) pushed from TradingView webhooks, keeps a rolling window per timeframe, and derives a regime for each sleeve (for example STRONG_BULLISH, BULLISH_PULLBACK, CHOPPY_SIDEWAYS).
  2. Score each bar with a decision model. Every bar close is compiled into a structured market snapshot and sent to TypeSafe Jev, which returns a typed verdict: an action, a setup grade, and an exhaustion-risk figure.
  3. Simulate execution honestly. A paper book applies real Bursa friction (1 tick = RM 1 per ton = RM 25 per contract, RM 15 round-turn commission, 1-tick slippage) so the record reflects what a real account would experience.

The name is deliberate. The interface and the decision loop are modelled on jev-trade, a crypto perpetuals trading tool. Jev FCPO ports that shape to Malaysian palm oil futures, swapping the crypto tick model for Bursa's contract spec and replacing the generative decision logic with TypeSafe Jev.

What problem this solves

A retail FCPO trader working a single chart has to hold four timeframes in their head at once: the 30-minute tide, the 5-minute structure, the 1-minute trigger, and the risk maths that ties them together. Under a live session, that is where discipline breaks. Jev FCPO externalises the whole loop. It does not tell you what to feel; it prints the verdict and logs it, so the trader can benchmark their own calls against an unemotional, repeatable process.

The four-step rulebook the engine enforces

The decision model is not asked to freeform. It scores against a fixed 4-step trend SOP drawn from Elder's Triple Screen, Raschke's EMA pullback, Dow structure, and Wilder ATR sizing:

StepTimeframeQuestion answered
1. Macro tide30mWhich direction is allowed? Longs only, shorts only, or flat.
2. Structure5mHas a pullback reached the value zone between EMA 9 and EMA 21 without breaking the swing pivot?
3. Trigger1m / 3mDid the fast bar confirm entry (rejection wick, engulfing, volume expansion)?
4. RiskallIs there a deterministic stop and at least 1:2 reward-to-risk?

Two invariants sit above the steps: never trade against the 30m tide, and never open a new position within ten minutes of a session close. Everything the model can say is bounded by those rules.

The Jev FCPO terminal in dark theme: header with balance, session badge and paper-simulator notice; four market sleeves showing regime and price; a 5-minute candlestick chart in green and red with gold, blue and purple EMA 9/21/50 overlays; a large price readout and HOLD verdict overlaid on the chart; the Jev call rail with probability bars; the live calls feed with per-call latency; and the paper account book.
Fig 1. The Jev FCPO terminal, dark theme. Left to right: four timeframe sleeves (30m macro tide, 5m structure, 3m, 1m trigger) each with its derived regime; the candlestick chart with EMA 9/21/50 overlays and a whole-number price scale; the Jev call rail showing the live verdict with its long/short/hold probability split; the live calls feed logging every evaluation and its latency; and the paper account book at the base. The header carries the virtual balance (RM 50,000), realised PnL, win rate, the session badge, and the standing paper-simulator notice.
The same Jev FCPO terminal in light theme, showing that the session badge, paper-simulator notice, help modal, sleeve labels and profit/loss figures all remain legible on a light background.
Fig 2. The same terminal in light theme. The interface shipped with hardcoded dark-only colours, so status pills, the help modal, sleeve labels and every profit/loss figure became unreadable on a light background - white text on white. They now read from theme-aware tokens, and every probe point measures at a contrast ratio of at least 5.17:1 in both themes. This figure is the evidence for that claim rather than a restatement of it.

Why we are pushing Jev here

The build exists to test a specific claim: that a typed, non-generative decision model can hold a repeatable intraday process better than a human under pressure. That is why the engine logs every verdict with its latency and cost, why the paper book models real friction instead of pretending fills are free, and why this report bothers to measure the model rather than assert that it works. The question is not "can Jev produce a signal". The question is whether its signal is stable, affordable, and disciplined enough to be worth acting on. The rest of this report answers that with data from a live session.

2. Why Jev matters to this build

110
Real Jev calls
15 + 95, 0 simulator fallbacks
427 ms
p50 latency
n=27 benchmark
$0.032
Per 1,000 calls
$0.0000318 each
9
Signals rejected
gate held every time
0
Trades executed
never satisfied the gate
0
Setups scoring ≥ 2
in 95 single-sleeve calls

Jev FCPO is not a charting toy. It ingests multi-timeframe FCPO bars and asks a decision model a structured question: what is the best immediate action, how good is the setup, and how exhausted is the move? Everything downstream (the paper book, the stop sizing, the R:R filter) hangs off that answer.

The original code shipped an offline keyword simulator (jev-fcpo-sim-v1) because no model key existed. The simulator scored grade by matching words in a text prompt, which is why early sessions reported Grade A on every single bar. That is not a decision engine; it is a lookup table wearing a lanyard.

The first live session replaced it with the real model.

3. Integration contract

Jev is reachable two ways, and the code prefers them in this order:

PriorityPathAuthEndpoint
1TypeSafe directTYPESAFE_API_KEYapi.typesafe.ai/v1/systemone
2Jev via OpenRouter (used here)OPENROUTER_API_KEYopenrouter.ai/api/alpha/decisions
3Offline simulatornonein-process
The discovery that shaped the integration ~typesafe/jev-latest is a decisions model, not a chat model. Posting it to /chat/completions returns a 400.
400: "~typesafe/jev-latest is a decisions model and cannot be used with the
chat/completions endpoint. Use the /api/alpha/decisions endpoint instead."

The decisions endpoint accepts the exact { model, state, questions } contract the existing code already built for TypeSafe, so the real model drops in without reshaping the schema. Each question (choice, score, noul) must carry an instructions field or the payload is rejected.

Response shape

{
  "model": "typesafe/jev-1.13-20260917",
  "answers": {
    "action":          { "choice": "ENTER_LONG", "probabilities": { "ENTER_LONG": 1, "ENTER_SHORT": 0, "HOLD": 0 }, "confidence": 1 },
    "setup_quality":   { "score": 2.84, "probabilities": { "2": 0.15, "3": 0.85 }, "confidence": 0.84 },
    "exhaustion_risk": { "noul": 0.22 }
  },
  "usage": { "input_tokens": 442, "output_tokens": 78, "cost": 0.000018564 },
  "provider": "TypeSafe",
  "id": "gen-dec-1790003587-YJla9sL0ZkB5NbiVD2cv"
}

4. Measurements

Latency sample combines 15 live FCPO bar-close evaluations from the first session with a 12-call controlled benchmark across four synthetic market states (aligned-bullish, aligned-bearish, chop, conflict), 3 repeats each. Full data: docs/jev-bench.json. The 95-call instrumented run in section 4.6 is measured separately and is not folded into this sample.

4.1 First session: latency distribution

0 3 6 9 12 12 6 5 0 2 1 1 200-400 400-600 600-800 800-1000 1000-1200 1200-1400 1400-1600 Latency (ms) Calls SPEC target "<50ms" met by 0 of 27 calls
Fig 3. Latency histogram (200 ms bins). 59% of calls land under 500 ms and 85% under 1 s, but the distribution is long-tailed: one call reached 1431 ms, 3× the median. The coefficient of variation is 0.54, so per-call latency is not predictable. At bar-close cadence this is tolerable; the SPEC's "<50 ms" figure is aspirational and was met by zero calls.
StatisticValue (ms)
n27
min291
p50 (median)427
mean555.0
p901069
p951287
max1431
stdev301.8
CV (stdev/mean)0.54

4.2 First session: decision filter funnel

15 market evaluations 15 5 actionable ENTER_* signals 5 0 passed Grade B+ gate 0 0 executed paper trades 0 Bar length proportional to count gate: actionConfidence ≥ 0.75 · setupScore ≥ 2 · exhaustionRisk < 0.40
Fig 4. First-session entry-filter funnel. Jev proposed a direction on 5 of 15 bars, but every proposal graded C (score ≈ 1), below the Grade B threshold. The filter rejected all of them, so the paper book stayed flat. This is the intended behaviour: the engine distinguishes a textbook setup from a counter-trend scalp.

4.3 First session: verdict mix

Action HOLD · 10 (66.7%) ENTER_SHORT · 5 (33.3%) Setup grade Grade F · 10 (66.7%) Grade C · 5 (33.3%) No Grade A or Grade B appeared at any point in the session. The simulator, by contrast, reported Grade A on every bar.
Fig 5. First-session verdicts. Real Jev's output is bimodal: two thirds of bars were unambiguous chop (Grade F, HOLD), the rest were marginal counter-trend scalps (Grade C). Nothing reached Grade B. The absence of Grade A is the model being honest about a morning of conflicting timeframes, and it is the clearest single measure of the difference between the real model and the keyword simulator it replaced.

4.4 Determinism and outcome stability

The controlled benchmark repeated each of four market states three times. Scores are continuous and near-stable; the action is not, in the one ambiguous regime.

0 1 2 3 setup_quality score (0 = F, 3 = A) aligned-bullish 2.59-2.61 ENTER_LONG ×3 aligned-bearish 2.13-2.21 ENTER_SHORT ×3 chop 0.00 (zero spread) HOLD ×3 conflict 1.14-1.19 HOLD ×1 ENTER_LONG ×2
Fig 6. Score range and action consistency per state (3 repeats each). Clear regimes are reproducible: aligned-bullish held ENTER_LONG three times with a score spread of just 0.02. The conflict state (bearish 30m tide versus a strong bullish 5m/1m trigger) is not reproducible: two runs returned ENTER_LONG and one returned HOLD on a near-identical prompt. Jev is a sampled model, so a signal that sits on a regime boundary can flip between runs.
Market stateScore minScore maxSpreadActions (of 3)
aligned-bullish2.592.610.02ENTER_LONG ×3
aligned-bearish2.132.210.08ENTER_SHORT ×3
chop0.000.000.00HOLD ×3
conflict1.141.190.05ENTER_LONG ×2, HOLD ×1
Operational consequence Because the action can flip on boundary states, the grade gate is doing real work. Requiring Grade B (score ≥ 2) keeps the engine out of exactly the states where the model is least stable. The conflict state scores 1.14-1.19, comfortably below the gate, so its instability never reaches the paper book.

4.5 Second session: moving off the laptop, and learning to measure

The night session moved the engine off the laptop tunnel and onto the host that already runs the other fleet services, behind nginx and TLS. This matters more than it sounds. A Cloudflare quick tunnel is bound to a process on a laptop: close the lid, lose the feed. The engine now runs as a supervised service that survives a reboot, and the webhook endpoint is a stable public URL the TradingView alerts can keep pointing at week after week.

The deployment also forced one security change. The webhook accepted any request that simply omitted the secret field - harmless on localhost, an open door on a public endpoint. Anyone could have posted fabricated bars, and every bar triggers a paid model call and can move the paper book. The secret is now required, not optional. Verified: a request with no secret returns 401, the same as one with a wrong secret.

The second change was measurement. Until this session, Jev's verdicts lived in a 50-entry in-memory ring buffer that emptied on every restart, so the obvious question - when Jev said this, what actually happened? - was unanswerable after the fact. The engine now appends every evaluation to a durable log and stores the answer next to the question:

GroupCaptured per evaluation
Identitydecision id, timestamp, bar timestamp, symbol, timeframe, session
Market at decision timeOHLC, volume, ATR, RSI, and all four sleeve regimes
Jev's answeraction, action confidence, full probability vector, setup score and grade, exhaustion risk
Provenancemodel version, generation id, provider, and whether the call was real or simulated
Performancelatency, input/output tokens, and the real per-call cost returned by the API
Gate outcomewhether the gate applied, whether it passed, and the exact failing reason
Outcomeexecuted side and fill price, plus a separate close record linked back to the opening decision
Why the last row is the point Calibrating a threshold requires knowing which verdicts preceded which outcomes. Without the link between a decision and its close, you have a log of opinions, not evidence.

4.6 The instrumented run, and what it revealed

A 95-evaluation run exercised the full path end to end. The caveat belongs before the numbers: this run was driven by replayed bars generated for the test, not by the live market, though all 95 evaluations were real model calls against a synthesised snapshot. It is an integration test with a real decision model, not market evidence, and it is archived separately from the live log so it cannot contaminate calibration data.

score 0 (Grade F) 63 score 1 (Grade C) 32 score 2 (Grade B) 0 score 3 (Grade A) 0 The gate requires score ≥ 2. The distribution never reached it in 95 evaluations. Every snapshot carried NO_DATA for the 1m, 3m and 30m sleeves - Jev judged a four-step setup from the 5m sleeve alone.
Fig 7. Setup-score distribution over 95 real model calls. The scores are pinned at 0 and 1. Nothing could have passed the entry gate regardless of any other input, because the setup-quality score never reached the required 2. This is not a defect in the model or the gate: it is what happens when a multi-timeframe rulebook is asked to grade a setup with only one of its four timeframes feeding it.
95 market evaluations 95 4 actionable ENTER_* signals 4 0 passed the gate 0 0 executed paper trades 0 Same gate, one fifth the signal rate reject reasons: confidence < 0.75 ×4 · setupScore < 2 ×4
Fig 8. Entry-filter funnel for the instrumented run. Four signals in 95 evaluations, and all four failed on both criteria independently. Jev proposed a short while grading its own setup F, at roughly half the confidence the gate requires.
MetricValue
Evaluations95
Real model calls (0 simulator fallbacks)95
ActionsHOLD 91, ENTER_SHORT 4
GradesF 63, C 32
Setup score0 (n=63), 1 (n=32) - never reached 2
Action confidencemin 0.39, p50 0.84, mean 0.83, max 1.00
Exhaustion risk0.14 to 0.23
Latencymin 293, p50 363, p90 481, p99 781 ms
Cost$0.00339 total, $0.0000357 per call
Gate outcome4 signals, 0 passed, 4 rejected
Trades executed0

The four entry signals were the interesting part, and none was a close call:

SignalConfidenceSetup scoreGradeExhaustion
ENTER_SHORT0.390F0.21
ENTER_SHORT0.480F0.23
ENTER_SHORT0.490F0.22
ENTER_SHORT0.430F0.23
The finding to carry into a live session Setup score never reached 2 in 95 evaluations, so nothing could have passed regardless of any other input. The cause is structural, not strategic: every snapshot carried NO_DATA for the 1m, 3m and 30m sleeves, so Jev was grading a four-step multi-timeframe setup from the 5-minute sleeve alone. The rulebook's first step is the 30-minute macro tide; with no 30-minute data, a textbook Grade A is unreachable by construction, and Jev declined to invent one. The first session, which had all four sleeves connected, produced a genuine spread including Grade C entries. Confirm the 30m sleeve is streaming before drawing any threshold conclusion from the output.

4.7 Ingestion and cost

1m 22 3m 8 5m 5 30m 1 Bar counts are proportional to timeframe. 30m had a single close (11:30) inside the session window.
Fig 9. Bars ingested per sleeve. All four sleeves confirmed live via TradingView Pine webhooks. The 30m sleeve only closed once during the window, which is why the macro tide read NO_DATA for most of the session and flipped to BEARISH_PULLBACK at 11:30. A single 30m bar is not a trend.
MetricValue
Cost per call$0.0000318
Cost per 1,000 calls$0.0318
Mean input tokens757
Mean output tokens96
Cost for a 6-hour session, 4 sleeves$0.0036
Success rate27/27 (100%)

The cost figure is the decisive one: a session costs well under a cent. There is no economic case for the offline simulator whenever a key is present.

5. Defects these sessions exposed and fixed

1

ATR risk basis. Stops were sized from whichever bar triggered the call (usually the 1m, ATR ≈ 3.7) when CONTEXT.md requires the 5m structure timeframe. Now reads the 5m sleeve's ATR, falling back to the trigger bar.

2

Swing pivot ignored. The spec's primary stop basis, the 5m swing pivot, was captured in state but never passed to the paper book. Stops are now placed 1 tick beyond the structural swing when that is wider than the ATR stop, preserving 1:2 R:R.

3

Fractional grade crash. Jev returns continuous scores (2.84). The code indexed a 4-element array with the raw float, yielding undefined then Grade F. Now rounded before grading. Without this fix the benchmark's 2.59-2.61 scores would all have graded F.

4

Webhook body validation. Alerts configured with a default TradingView condition send their own text body, not the Pine JSON. The server now logs the raw body so misconfiguration is diagnosable rather than silent.

5

Optional webhook secret. The secret check only ran when a secret was present, so a request that simply omitted the field bypassed it entirely. Harmless on localhost; on a public endpoint anyone could have posted fabricated bars, and every bar triggers a paid model call and can move the paper book. Now mandatory. Found while preparing the deployment, not by a failure.

6

Decision log lost on restart. Verdicts were held in a 50-entry in-memory buffer and vanished when the process restarted, making it impossible to ask later what Jev said before a given outcome. Replaced with an append-only durable log that also records the gate result and the eventual close.

7

Theme tokens were hardcoded. Large parts of the interface were built for a dark background only: status pills, the help modal, sleeve labels, and every profit/loss figure. They did not merely look wrong in light theme, they became unreadable - white text on white. All of it now reads from theme-aware tokens, and every probe point measures at a contrast ratio of at least 5.17:1 in both themes.

8

Chart signals were invisible. The chart plotted filled trades but not Jev's rejected proposals, so on a session where nothing passed the gate the chart appeared to do nothing at all. Rejected signals now render as distinct markers, separated by timeframe so a 30m alert cannot drop a marker into the 1m chart.

6. Conclusions and carry-forward

  1. Jev is a decisions model, not a language model. It refuses chat endpoints and speaks a typed choice/score/noul contract. That contract already existed in the codebase; the integration was a transport swap.
  2. Latency is hundreds of milliseconds, not tens. p50 427 ms, p95 1287 ms, CV 0.54. Fine for bar-close decisions, not for anything tick-driven.
  3. Clear regimes are reproducible; boundary regimes are not. Aligned states repeated identically across runs; the one conflicting state flipped action on an identical prompt. The grade gate is the guard that keeps this out of the paper book.
  4. Cheap enough to run unconditionally. $0.0318 per 1,000 calls removes any cost argument for the simulator.
  5. Decision quality is bounded by how many sleeves are feeding it. With all four connected, the model produced a genuine spread including Grade C entries. With only the 5m sleeve, the setup score never left the bottom two rungs in 95 calls. Check ingestion before interpreting any threshold result.
  6. Require the webhook secret. An optional check is not a check. This is the difference between a localhost toy and a public endpoint.
  7. Measure before tuning. Every threshold decision in the next phase depends on the durable log, which did not exist before this work.
  8. Operational traps to remember. Do not run the simulation script against a live engine - test_rich_feed.mjs pushes ~4200-price bars into the same in-memory state as live ~4905 bars and fabricates positions. Validate timeframes on ingest, since anything outside 1/3/5/30 is accepted but invisible to /api/status and the UI.
What is proven, and what is not Proven: the plumbing, the contract, the measured envelope, the durable measurement path, and the filter's ability to reject marginal signals even when the model itself is uncertain. Not proven: any edge. Two sessions of real market data (110 model calls) produced nine entry signals and zero trades. That zero has two different causes, and conflating them would be wrong. In the first session, with all four sleeves streaming, Jev simply never graded a live bar above C, so the filter had nothing worth passing. In the second, single-sleeve run, the setup score never reached the gate's threshold at all. Neither is evidence that the gate is too strict: the controlled benchmark shows the model does grade a clean aligned state at score 3 (Grade A) with 0.92 confidence, and that combination passes the gate comfortably. The engine has not yet taken a trade because it has not yet been shown a live setup of that quality, not because the thresholds are unreachable. Threshold calibration against outcomes remains the next task, and it cannot start until a session runs with all four sleeves streaming and a full bar history accumulated.