Time-Series Validation: Preventing Leakage from Late Data and Future Labels

Build time-series validation around feature availability and label cutoffs. A reproducible Python experiment exposes leakage that chronological splits cannot catch.

13 min read

A forecasting model can pass a chronological train/test split and still use information that did not exist when its predictions were supposed to be made. A delayed measurement, a revised historical record, or a target-derived feature can cross that boundary without moving a single row between the training and test sets.

Consider an hourly demand forecast issued at 12:00 for demand at 18:00. Measurements arrive two hours after their event timestamps. At 12:00, the latest usable observation is the one timestamped 10:00. A feature called lag_1 sounds historical, but the 11:00 measurement will not arrive until 13:00.

This constructed example gives us a precise question: could the same feature vector and fitted model have existed at the recorded decision time? We will turn that question into executable checks, then run a small forecasting experiment with both valid features and a deliberately invalid positive control.

Put availability time into the data contract

Three timestamps serve different purposes. Event time says when something happened. Availability time says when the prediction system could use it. Decision time says when a particular forecast was issued. A historical database usually makes event time easy to find; availability and revision history need deliberate preservation.

Item Event time Available to this system
Latest eligible observation at 12:00 10:00 12:00
Tempting but unavailable lag-one value 11:00 13:00
Target of the 12:00 forecast 18:00 20:00

For this laboratory, arrivals exactly at a cutoff are processed before fitting or prediction. Thus available_at <= decision_at is eligible. If your scheduler runs before ingestion at the same timestamp, use a stricter boundary or a sequence number. Equality is an operational choice that deserves a test.

The two-hour delay is fixed, observations are immutable, timestamps are timezone-aware, and the series has an uninterrupted hourly grid. These assumptions let us represent availability with a shift. They are not a claim that a two-row shift works for irregular arrivals or revised records.

For a production source, record entity identity, event time, ingestion or serving availability, revision identity, and value. Availability should include the stages required before the model can actually read the value. Landing in object storage is insufficient if the feature materialization job has not finished.

A chronological split protects only one boundary

TimeSeriesSplit builds ordered training and validation partitions. Its documentation also calls out equal spacing when comparing folds covering equal durations. The splitter operates on row positions; it does not inspect how your feature columns were produced or when labels became available.

There are at least three independent checks. Each validation feature must be available at its own prediction time. Each training label must be available at the model’s fitting cutoff. Learned preprocessing must be fitted using eligible training rows. Passing one does not establish the others.

Random splitting can also answer the wrong deployment question by mixing later regimes into training. The scikit-learn lagged-feature example demonstrates an optimistic shuffled evaluation on its dataset. That example does not imply every shuffled score must improve. Here we keep the split chronological throughout, so the experiment isolates feature availability.

Build features from the history that has arrived

The tested environment is Python 3.13.4, NumPy 2.5.3, pandas 3.0.6, and scikit-learn 1.9.1. Save the following as temporal.py. It contains the feature builder, label eligibility filter, model factory, and a mutation-based audit.

"""A fixed-delay, single-series laboratory; not a general feature store."""
import numpy as np
import pandas as pd
from sklearn.linear_model import Ridge
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

HORIZON = 6
DELAY = 2
FEATURES = ["latest", "mean_24", "previous_day", "hour_sin", "hour_cos"]


def validate_series(y):
    index = y.index
    if not isinstance(index, pd.DatetimeIndex) or index.tz is None:
        raise ValueError("a timezone-aware DatetimeIndex is required")
    if len(y) < 40 or not index.is_unique or not index.is_monotonic_increasing:
        raise ValueError("need at least 40 unique, ordered observations")
    if not (index.to_series().diff().iloc[1:] == pd.Timedelta(hours=1)).all():
        raise ValueError("this laboratory requires an uninterrupted hourly grid")
    if not np.isfinite(y.to_numpy(dtype=float)).all():
        raise ValueError("missing or nonfinite observations need an explicit policy")


def features(y):
    validate_series(y)
    latest = y.shift(DELAY)
    target_hour = (y.index.tz_convert("UTC") + pd.Timedelta(hours=HORIZON)).hour
    return pd.DataFrame({
        "latest": latest,
        "mean_24": latest.rolling(24, min_periods=24).mean(),
        "previous_day": latest.shift(24),
        "hour_sin": np.sin(2 * np.pi * target_hour / 24),
        "hour_cos": np.cos(2 * np.pi * target_hour / 24),
    }, index=y.index)


def dataset(y):
    frame = features(y)
    frame["target"] = y.shift(-HORIZON)
    frame["label_available_at"] = y.index + pd.Timedelta(hours=HORIZON + DELAY)
    # Drop only the deterministic warm-up and unavailable target tail.
    return frame.dropna()


def eligible_training_rows(frame, candidates, fit_at):
    rows = frame.iloc[candidates]
    keep = (rows.index < fit_at) & (rows["label_available_at"] <= fit_at)
    selected = np.asarray(candidates)[np.asarray(keep)]
    if len(selected) == 0:
        raise ValueError("no labels were available at the fit cutoff")
    return selected


def model():
    return make_pipeline(StandardScaler(), Ridge(alpha=1.0))


def assert_available_history_only(builder, y, decision_at):
    """A counterexample test, not a proof for every possible input."""
    changed = y.copy()
    unavailable = y.index + pd.Timedelta(hours=DELAY) > decision_at
    changed.loc[unavailable] += 10_000
    before = builder(y).loc[decision_at]
    after = builder(changed).loc[decision_at]
    np.testing.assert_allclose(before, after, rtol=0, atol=1e-9)

At decision time t, latest is the observation at t - 2 hours. The trailing mean contains 24 observations ending there. previous_day shifts that already-delayed series another 24 rows, so its source is t - 26 hours. The daily sine and cosine describe the target hour in UTC, which is known when the forecast is issued.

Positive shift periods move values into later rows; we do not pass freq, which changes timestamp labels rather than producing this value alignment. See the pandas shift reference. The grid validation matters because two rows represent two hours only under the stated contract.

A trailing rolling window still needs that availability shift. A centered window is an additional hazard because it can include observations after the row’s timestamp. The rolling-window reference describes the window alignment and minimum-observation controls. Here the first incomplete windows remain missing; we do not backfill them with later values.

Computing these particular lag features over the complete historical array is safe under the fixed-delay, immutable-data assumptions: each row reads only its permitted prefix, and the transformation learns no population statistics. The same argument would fail for a full-series normalization, global feature selection, or a rolling calculation with future observations.

Training labels need their own cutoff

A row whose forecast origin is 12:00 has a target event at 18:00 and an available label at 20:00. It cannot be used to train a model fitted at 19:00, even though the row timestamp and target event both precede that fitting time.

At a validation start v, the eligibility condition in this example is training_origin + 6 hours + 2 hours <= v. The latest eligible training origin is therefore v - 8 hours. An ordinary ordered split would include origins through v - 1 hour; seven rows must be removed.

This is why gap=7 is equivalent to our timestamp filter on this exact grid and cutoff convention. The last retained training position lies eight hourly steps before the first validation position. The test suite checks this equivalence rather than relying on an intuitive guess that the gap should equal the forecast horizon.

Keep the timestamp rule as the source of truth. Variable label delays, missing hours, different target horizons, or a fitting job that starts earlier invalidate that simple row-count shortcut. With a real fitting duration, use the input snapshot cutoff that precedes model availability; this laboratory treats fitting at the boundary as instantaneous.

A gap chosen for label availability also does not make residuals independent or eliminate every overlap problem. Targets defined as sums over future intervals need their full outcome interval and finalization delay recorded. Validation design must reflect how those outcomes become observable.

Keep learned preprocessing inside each fit

The model factory returns a fresh StandardScaler and Ridge pipeline for every fold and feature variant. The scaler learns its mean and scale from the eligible training rows. Validation rows receive the resulting transformation without changing those fitted statistics.

This follows scikit-learn’s guidance on preprocessing leakage. The same training-only rule applies to learned imputation values, dimensionality reduction, feature selection, and target encodings. The Pipeline API provides a way to fit and apply those steps together.

A pipeline cannot identify a forbidden input column by its meaning. If a column contains tomorrow’s outcome, fitting its scaler on training data correctly does not make that value available when today’s forecast is issued. We deliberately include such a column next to prove that a valid split and a valid pipeline can coexist with invalid features.

Run a controlled leakage experiment

Download the complete laboratory, dependency lock, and temporal-boundary tests (ZIP). The package includes a runner that records versions, output, and test results. It uses synthetic data and downloads no dataset.

For the two inline files, create an isolated environment and install the tested dependencies:

python3 -m venv .venv
source .venv/bin/activate
python -m pip install numpy==2.5.3 pandas==3.0.6 scikit-learn==1.9.1

Save this second file as demo.py beside temporal.py, then run python demo.py.

import numpy as np
import pandas as pd
from sklearn.metrics import mean_absolute_error
from sklearn.model_selection import TimeSeriesSplit
from temporal import FEATURES, dataset, eligible_training_rows, model


def synthetic_series(n=2400):
    rng = np.random.default_rng(20260918)
    hours = np.arange(n)
    residual = np.zeros(n)
    noise = rng.normal(0, 5, n)
    for i in range(1, n):
        residual[i] = 0.85 * residual[i - 1] + noise[i]
    values = 100 + 20 * np.sin(2 * np.pi * hours / 24) + residual
    return pd.Series(values, index=pd.date_range(
        "2025-01-01", periods=n, freq="h", tz="UTC"), name="demand")


def evaluate():
    frame = dataset(synthetic_series())
    # Deliberately invalid positive control: the outcome is supplied as input.
    frame["future_value"] = frame["target"]
    splitter = TimeSeriesSplit(n_splits=4, test_size=240)
    results = []
    for fold, (candidates, test) in enumerate(splitter.split(frame), 1):
        fit_at = frame.index[test[0]]
        train = eligible_training_rows(frame, candidates, fit_at)
        row = {"fold": fold, "removed": len(candidates) - len(train)}
        for name, columns in [("valid", FEATURES),
                              ("leaked", FEATURES + ["future_value"])]:
            fitted = model().fit(frame.iloc[train][columns], frame.iloc[train]["target"])
            predicted = fitted.predict(frame.iloc[test][columns])
            row[name] = mean_absolute_error(frame.iloc[test]["target"], predicted)
        row["baseline"] = mean_absolute_error(
            frame.iloc[test]["target"], frame.iloc[test]["latest"])
        results.append(row)
    return results


if __name__ == "__main__":
    print("fold  purged  latest-value MAE  valid MAE  leaked MAE")
    for row in evaluate():
        print(f"{row['fold']:4d}  {row['removed']:6d}  {row['baseline']:16.3f}"
              f"  {row['valid']:9.3f}  {row['leaked']:10.3f}")

The generated series combines a daily cycle with correlated noise. Each of four validation folds contains 240 hourly forecast origins. The model is fitted once at the start of a fold; features then advance with each origin as new observations become available. This is a sequence of six-hour-ahead forecasts using a fixed model within each fold.

It is not a single forecast issued at the fold boundary for all 240 future hours. That protocol could not use observations arriving later in the fold. Matching this distinction to your serving schedule is as important as selecting the splitter.

Mean absolute error (MAE) averages the absolute difference between predictions and observed targets, in the synthetic demand units used here. Scores are computed retrospectively after the validation labels have arrived. Label availability during model fitting is enforced separately.

The actual output from the pinned environment was:

fold  purged  latest-value MAE  valid MAE  leaked MAE
   1       7            22.597      6.811       0.016
   2       7            21.940      8.427       0.017
   3       7            23.388      6.845       0.012
   4       7            22.034      6.507       0.010

The latest-value baseline predicts the most recently available observation for every horizon. The valid model also receives the trailing mean, previous-day value, and calendar terms. The leaked variant receives the actual future target as an extra feature. Its tiny error is an intentionally invalid positive control, not a useful forecasting improvement.

Both model variants use the same validation origins, label cutoff, and Ridge settings. That controlled comparison demonstrates a concrete failure: chronological partitioning and training-only scaling do not remove information already embedded in the feature matrix. It does not estimate leakage prevalence or performance on a real demand dataset.

Turn temporal assumptions into regression tests

The mutation audit changes every observation unavailable at a selected decision time, including recent observations that happened earlier but have not arrived yet. It then rebuilds the feature vector at that same decision time. A valid history-only builder should produce the same vector.

Our tests apply that audit at multiple cutoffs. They also verify that it rejects both a direct future-target feature and the more plausible one-row lag under the two-hour reporting delay. Those deliberately bad builders confirm that the audit can detect the failures it was designed to catch.

A passing mutation test is evidence for the selected inputs and perturbations, not a mathematical proof. A transformation that clips values or branches on special cases may require different perturbations. Combine this test with a review of source dependencies, timestamp joins, and training snapshots.

The 15-test suite also checks exact feature values, target alignment, arrival-at-cutoff equality, individually delayed labels, scaler statistics, rejected irregular or naive timestamps, nonfinite inputs, timezone conversion, and removal of the expected warm-up and target tail. Run python run_lab.py from the extracted package to reproduce the experiment and all tests.

Revisions and joins are where simple shifts stop working

A real historical table may contain the final corrected value for every date. Replaying that table can leak corrections that arrived days after the original prediction. Shifting final values backward by a nominal delay does not reconstruct what the service actually knew.

Point-in-time retrieval is the relevant family of techniques; Feast’s point-in-time join documentation illustrates historical feature retrieval at entity timestamps. For your system, inspect what the stored timestamps mean and how late arrivals and revisions are represented. A join on event time alone cannot enforce an availability constraint that was never stored.

With revision history, select only versions available by the decision cutoff, apply the intended event-time window, and reproduce the revision selection rule used by serving. Preserve entity boundaries too: a lag across two concatenated customers is not a customer’s history. The single-series code here deliberately does not implement those joins.

Known future inputs require a different distinction. A published holiday calendar can be valid; a realized future weather measurement cannot stand in for the weather forecast available at the time. Retain forecast vintages or other historical input snapshots when that is what production reads.

Keep the final evaluation outside model selection

The four folds here are a diagnostic laboratory with fixed hyperparameters. If you use fold scores to choose lags, penalties, features, or retraining schedules, those folds become development data. Reserve a later untouched period for the final evaluation and freeze the selection process before examining it.

Report results by horizon and relevant operational slices, alongside a baseline that obeys the same information cutoff. Do not treat adjacent hourly errors as independent observations when building uncertainty estimates. A single average also hides periods with delayed feeds or unusual demand.

The practical deliverable is an auditable timeline: the feature snapshot used for each decision, the labels eligible at each fit, the fitted preprocessing state, and the resulting predictions. When those artifacts agree with the serving contract, the score describes a forecasting task the system could actually have performed.

What do you think?

Add your perspective.

Your email address will not be published. Required fields are marked *

What is on your mind?

START WITH A TOPIC
SAVED FOR A QUIET MOMENT

My reading list

Your list is stored only in this browser.

See you in the next story.

New ideas, new stories. The same curiosity.

Open the RSS feed