A useful comparison
starts with a clear method.
What happened in comparable situations?
The starting point is a subject: a symbol, a date, and a chart timeframe. The research asks which historical observations provide a defensible comparison, what followed those observations, and how well that distribution describes new cases.
These are separate questions. A strong visual match does not establish a causal explanation, a directional signal, or a profitable strategy.
Describe the state before examining its outcome.
The retrieval system represents historical chart windows using learned embeddings: numerical descriptions that let the system compare their structure. The shape representation is trained without forward-return labels. Retrieval similarity and subsequent returns have different roles.
Chart structure alone is not the whole situation. Market conditions and events can affect which comparisons are meaningful. The public studies test whether proposed context filters or ranking changes improve a defined outcome measure; context is not promoted merely because its story sounds plausible.
When reading a response, inspect the actual selected sample and any condition or event-overlap information provided. Do not assume every retrieved member shares the subject's event.
Show the distribution, not just its center.
After a historical set is selected, the system summarizes what happened next at specified horizons. A median describes the middle of those observations. Percentiles describe their spread. An up rate describes a frequency in that sample.
None of these is an individualized probability guarantee. Horizons, return definitions, sample sizes, and data dates matter. Some research reports excess returns relative to the market; other tools report raw returns. Read the definition attached to the output.
Ask whether the ranges actually hold.
A nominal 80% interval aims to contain 80% of outcomes under the relevant evaluation protocol. Empirical coverage is the fraction it actually contained. That distinction is the reason for a calibration record.
For its calibrated cohort-band method, Chart Library uses corrections to address under-coverage. Evaluation must consider both coverage and width: widening a band may improve coverage while making it less informative.
A fixed held-out evaluation answers a historical question about a method. A rolling service audit asks how published bands behave as new observations settle. Neither should be substituted for the other, and aggregate coverage can hide failures in a particular condition.
The calibration receipt applies only to the cohort-band method and population named in that receipt. Its coverage percentage does not automatically apply to the raw market-state excess-return ranges shown in the stock comparison tool. Those ranges are empirical historical observations, not automatically calibrated forecasts.
Daily research notes have their own settled-note tally. A single-horizon interval is not a guarantee that an entire multi-session path stays inside it.
Make the decision criteria explicit.
The public ledger pairs research questions with their stated decision criteria and available specification and result documents. Readers can examine the evaluated population, test protocol, and reason for a pass, fail, or other status.
For pre-registered work, defining the population and test before opening outcomes helps distinguish a genuine test from an explanation fitted after the fact. A failed sample should not quietly become a fresh confirmation sample after the method is adjusted.
New research can build on a negative finding, but the new question, changed method, and evaluation sample need to be distinguishable from the original. Study 102's event-day under-coverage and study 103's separately fitted gap-day correction illustrate that progression.
The ledger is a selected public collection, not a claim that every internal experiment is represented.
Where the evidence ends.
- Historical selection is not random sampling. Retrieved analogs can be dependent, sparse, or unrepresentative of a new situation.
- Time matters. Data coverage, available disclosures, market regimes, and the served method can change. Check the as-of date and evaluated period.
- Context can fail to help. More detailed conditions may reduce sample size without improving the result. Read the negative studies.
- Coverage is not trading performance. Costs, execution, liquidity, and portfolio risk are separate questions.
- A weak result may support no conclusion. Missing data, a failed coverage check, or poor comparability can be reasons to abstain.
Follow the evidence.
The research ledger links study specifications and results. Frozen evaluations preserve dated method evidence. The live coverage record exposes current audit status. The API documentation describes the structured research tools.
Source documents may refer to internal files or licensed input data that are not publicly downloadable. Open access to the research and tools does not imply unrestricted redistribution rights over every underlying dataset. See data and licensing and the terms of use.