Skip to content
Research
8 min read

The half of Borsa İstanbul that is missing from your backtest

927 companies have traded on Borsa İstanbul since 2000 and 612 are listed today. Build a universe from today's tickers and 315 firms, every one of them a company that failed, quietly vanish. Here is how large that error is, window by window, and the public archive that removes it.

Universe
927 tickers, 134 index reviews, 2000–2026
Test period
2000–2026

927 companies have traded on Borsa İstanbul since 2000. 612 are listed today.

Build a backtest universe from today's listings, which is the obvious thing to do and what most data tooling hands you by default, and you have silently deleted 315 companies. Every firm that went bust, was taken private, merged away, or was thrown off the exchange. The backtest that comes out the other side is measuring a world in which failure did not happen.

It will look good. It will not be true.

This article is about how large that error is on one specific market, and about the archive we put together to remove it, which is public.

How large the error is

The dataset behind this article is every BIST index-review snapshot from 2000-01-04 to 2026-10-01: 134 review dates, 53,983 rows, one per company per snapshot. Each row says what was listed on that date, which is the only thing a point-in-time universe needs to know.

Listed on that date612.00…of those, gone today0.00
0.00200.00400.00600.00200020012002200320042005200620072008200920102011201220132014201520162017201820192020202120222023202420252026
Fig. 1For each review date: how many companies were listed then, and how many of those no longer exist. The gap between the lines is what a universe built from today's tickers throws away.

Read the left edge first. In 2000, 329 companies were listed and 175 of them, 53.2%, are not on the exchange now. A study of "the Turkish market since 2000" built from today's tickers is a study of the half that made it.

The orange line falls as it approaches the present, and that is not good news. It falls because recent delistings have not happened yet. A window that ends today always looks the cleanest it will ever look.

WindowCompanies tradeableStill listed todayGoneShare gone
2005–201037220117146%
2010–201550731319438.3%
2015–202047634912726.7%
2020–2026672612608.9%

The last column is the size of the bias, window by window. Backtest a strategy over 2010–2015 against today's listed companies and roughly 38.3% of the names that existed are missing, and they are not a random 38.3%. They are the ones that failed.

Delistings are not a steady trickle

The other thing a single average hides is that companies leave in waves.

PeriodJoinedLeft
2001–20054651
2006–20108253
2011–201516377
2016–20205076
2021–202523152

2016–2020 is the only block where more companies left than joined, and 2021–2025 is the listing boom that followed. Both are periods a backtest badly wants to get right, and both are periods where a survivorship-contaminated universe gets them wrong in a specific direction: it removes the casualties from the bad stretch and keeps every winner from the good one.

Two traps inside the data

Both are properties of the source rather than bugs, and both cost us time.

The index column is a tier, not a membership. The source workbook has one cell per company per date, so it stores a single exclusive label: XU030 for the top thirty, XU050 for ranks 31–50, XU100 for ranks 51–100. They partition the BIST 100: thirty plus twenty plus fifty. Filter the raw column for XU050 expecting the BIST 50 and you get twenty companies, quietly, with no error anywhere.

Review dates are published before the reviews. BIST publishes the schedule of index reviews ahead of the reviews themselves, so the most recent snapshot can list every company while saying nothing about which index any of them is in. Ask it for the BIST 100 and a naïve implementation hands back an empty set: technically true, entirely misleading. In the library that case raises instead, and names the most recent review that was actually published.

Using it

from bist_constituents import constituents_on, tickers_in_range, Index

# The market as it actually stood on a given day
market = constituents_on("2014-06-30")
len(market)                                    # 446

# The BIST 30 that day: the real one, not today's
constituents_on("2014-06-30", index=Index.XU030).tickers

# A backtest universe: everything tradeable during the window
tickers_in_range("2015-01-01", "2020-01-01")   # 476 tickers

tickers_in_range is the one that matters for a backtest. It does not forward-fill across the leading edge: a company delisted before the window starts never appears, and one that died halfway through does. What comes back is the set of names genuinely available to trade in that window, and no others.

The repository ships the data as parquet and as CSV, so nothing forces you through the library at all:

import pandas as pd
df = pd.read_parquet("data/bist_constituents.parquet")

What this data does not tell you

  • No prices, volumes or fundamentals. Membership only. Pair it with a price source; the tickers are BIST codes and join directly.
  • No delisting reason. The data records that a company stopped trading, not whether that was bankruptcy, a merger or a voluntary exit, which matters a great deal if you are modelling what a holder actually recovered.
  • Quarterly, not daily. Membership changes on review dates. A company that listed and died between two reviews may not appear at all, so this removes most of the survivorship error rather than all of it.
  • Names are as recorded then. Companies rename, and the same ticker can carry different names across snapshots. That is the point of a point-in-time archive, and it means joining on name rather than ticker will bite you.

The raw workbooks are published by Borsa İstanbul and are not redistributed in the repository; the parser that turns them into this dataset is, so the archive can be rebuilt from source rather than taken on trust.

Every BIST backtest published on this site is built on this universe. That is the reason it exists. The reason it is public is that a result nobody can check the universe of is not worth much.

Written by

Koray GocmenFounder

Builds and runs the systematic strategies behind this research.

Full source

The code, the data and the parameter files behind this article.

All research