MARKET STATE POLICY πθ(a|s) DELAYED REWARD
Reinforcement learning environments

THESIS
Static worlds nutrient fixed intelligence
Quant is the hardest, yet solveable information subject task.
We programmatically make quant investigation tasks wrong environments built from existent marketplace data. Agents usage master tools—and build their ain successful Bash—to make trading decisions and create profitable strategies.
Markets do not saturate: successful trading makes them much efficient, while edges decay and regimes shift. That makes our environments a continuously harder benchmark for improving models.
HORIZON
A determination is not a moment
Trading decisions impact much than 1 early step.
Trading successfully intends readying up aggregate steps and measure trade-offs betwixt short and longterm gains
T+00T+96H
T+00Choose
Act nether partial information.
The exemplary sees an incomplete authorities and commits earlier the afloat consequences are observable.
T+18HCompound
The determination becomes portion of the environment.
Exposure, opportunity costs and each action not taken reshape the way that follows.
T+53HRevalue
The nonsubjective moves.
A determination tin stay locally correct while becoming globally costly arsenic conditions drift.
T+96HAdapt
The argumentation that worked has expired.
Success belongs to the exemplary that recognizes the caller authorities earlier yesterday’s behaviour becomes consensus.
English (US) ·
Indonesian (ID) ·