927 companies have traded on Borsa İstanbul since 2000. 612 are listed today.
Build a backtest universe from today's listings, which is the obvious thing to do and what most data tooling hands you by default, and you have silently deleted 315 companies. Every firm that went bust, was taken private, merged away, or was thrown off the exchange. The backtest that comes out the other side is measuring a world in which failure did not happen.
It will look good. It will not be true.
This article is about how large that error is on one specific market, and about the archive we put together to remove it, which is public.
How large the error is
The dataset behind this article is every BIST index-review snapshot from 2000-01-04 to 2026-10-01: 134 review dates, 53,983 rows, one per company per snapshot. Each row says what was listed on that date, which is the only thing a point-in-time universe needs to know.
Read the left edge first. In 2000, 329 companies were listed and 175 of them, 53.2%, are not on the exchange now. A study of "the Turkish market since 2000" built from today's tickers is a study of the half that made it.
The orange line falls as it approaches the present, and that is not good news. It falls because recent delistings have not happened yet. A window that ends today always looks the cleanest it will ever look.
| Window | Companies tradeable | Still listed today | Gone | Share gone |
|---|---|---|---|---|
| 2005–2010 | 372 | 201 | 171 | 46% |
| 2010–2015 | 507 | 313 | 194 | 38.3% |
| 2015–2020 | 476 | 349 | 127 | 26.7% |
| 2020–2026 | 672 | 612 | 60 | 8.9% |
The last column is the size of the bias, window by window. Backtest a strategy over 2010–2015 against today's listed companies and roughly 38.3% of the names that existed are missing, and they are not a random 38.3%. They are the ones that failed.
Delistings are not a steady trickle
The other thing a single average hides is that companies leave in waves.
| Period | Joined | Left |
|---|---|---|
| 2001–2005 | 46 | 51 |
| 2006–2010 | 82 | 53 |
| 2011–2015 | 163 | 77 |
| 2016–2020 | 50 | 76 |
| 2021–2025 | 231 | 52 |
2016–2020 is the only block where more companies left than joined, and 2021–2025 is the listing boom that followed. Both are periods a backtest badly wants to get right, and both are periods where a survivorship-contaminated universe gets them wrong in a specific direction: it removes the casualties from the bad stretch and keeps every winner from the good one.
Two traps inside the data
Both are properties of the source rather than bugs, and both cost us time.
The index column is a tier, not a membership. The source workbook has one
cell per company per date, so it stores a single exclusive label: XU030 for
the top thirty, XU050 for ranks 31–50, XU100 for ranks 51–100. They
partition the BIST 100: thirty plus twenty plus fifty. Filter the raw column
for XU050 expecting the BIST 50 and you get twenty companies, quietly, with
no error anywhere.
Review dates are published before the reviews. BIST publishes the schedule of index reviews ahead of the reviews themselves, so the most recent snapshot can list every company while saying nothing about which index any of them is in. Ask it for the BIST 100 and a naïve implementation hands back an empty set: technically true, entirely misleading. In the library that case raises instead, and names the most recent review that was actually published.
Using it
from bist_constituents import constituents_on, tickers_in_range, Index
# The market as it actually stood on a given day
market = constituents_on("2014-06-30")
len(market) # 446
# The BIST 30 that day: the real one, not today's
constituents_on("2014-06-30", index=Index.XU030).tickers
# A backtest universe: everything tradeable during the window
tickers_in_range("2015-01-01", "2020-01-01") # 476 tickerstickers_in_range is the one that matters for a backtest. It does not
forward-fill across the leading edge: a company delisted before the window
starts never appears, and one that died halfway through does. What comes back is
the set of names genuinely available to trade in that window, and no others.
The repository ships the data as parquet and as CSV, so nothing forces you through the library at all:
import pandas as pd
df = pd.read_parquet("data/bist_constituents.parquet")What this data does not tell you
- No prices, volumes or fundamentals. Membership only. Pair it with a price source; the tickers are BIST codes and join directly.
- No delisting reason. The data records that a company stopped trading, not whether that was bankruptcy, a merger or a voluntary exit, which matters a great deal if you are modelling what a holder actually recovered.
- Quarterly, not daily. Membership changes on review dates. A company that listed and died between two reviews may not appear at all, so this removes most of the survivorship error rather than all of it.
- Names are as recorded then. Companies rename, and the same ticker can carry different names across snapshots. That is the point of a point-in-time archive, and it means joining on name rather than ticker will bite you.
The raw workbooks are published by Borsa İstanbul and are not redistributed in the repository; the parser that turns them into this dataset is, so the archive can be rebuilt from source rather than taken on trust.
Every BIST backtest published on this site is built on this universe. That is the reason it exists. The reason it is public is that a result nobody can check the universe of is not worth much.