Skip to main content
Developer & API

The BankRegReports Data API: Complete Guide and Origin Story

The first time you sit down to pull bank financial data for a real project, you run into a wall that nobody warns you about.

14 MIN READ

The first time you sit down to pull bank financial data for a real project, you run into a wall that nobody warns you about.

The data is technically public. The FDIC posts quarterly call report files. The FFIEC runs a bulk data portal. The Federal Reserve publishes NIC data on every bank holding company. The NCUA has 5300 call reports for every credit union in the country. On paper, everything you need is out there, free, waiting.

Then you download it.

The FFIEC bulk files come as tab-delimited text, one file per schedule each quarter, with field names that are MDRM codes, strings like RCON2170 and RIAD4107 that mean nothing without a 600-page data dictionary. The FDIC files use a completely different identifier system. The Federal Reserve uses a third. None of them link to each other. Values for the same metric (say, total assets) show up in different units depending on the source, the quarter, and sometimes the specific schedule. FFIEC changed the way it reports certain interest income fields three times between 2019 and 2023 without publishing a migration guide.

You spend a week cleaning. Then another week figuring out why your NIM calculation does not match the UBPR. Then you discover that the UBPR itself is a derived product that applies specific peer-group adjustments, and that replicating it from raw call report data requires understanding which of the 3,819 live MDRM codes map to which UBPR line items, and that mapping is not published anywhere in machine-readable form.

At that point, you start looking at what the enterprise data platforms charge. The numbers are not friendly. Multi-year contracts. Seat licenses. Usage-based overages. The platforms built for institutional buy-side teams are priced for institutional buy-side budgets.

That gap, between “technically free but operationally unusable” and “clean but priced for a Bloomberg terminal,” is what BankRegReports exists to close.


Why Did We Build This?

We came to this project from the practitioner side. ALCO work. Regulatory exam preparation. Credit committee analysis. The kind of analysis where you need accurate, cross-source bank data and you need it fast, not six weeks of data engineering first.

We had three specific frustrations that we kept running into:

The calculation problem. The efficiency ratio your core banking system reports and the efficiency ratio your examiner uses are often different numbers. Not because either is wrong, but because there are four common definitions in active use across the industry and the formula varies by who’s asking the question. The same is true for NIM, for cost of deposits, for the allowance coverage ratio. If you pull raw numbers and compute these yourself without knowing which formula chain to use, you get a number that will not reconcile against regulatory benchmarks. We reverse-engineered the UBPR formulas, cross-validated against FFIEC ground truth across about 4,300 banks and 20 quarters, and baked the correct methodology into every metric we serve.

The linkage problem. A bank holding company files a Y-9C with the Federal Reserve. Its subsidiary banks file call reports with their primary regulators. The holding company has SEC filings. The whole structure is documented in NIC data. But the RSSD ID used in NIC does not map cleanly to the FDIC certificate number or the FFIEC reporter ID without a crosswalk that nobody publishes in a usable form. We built that crosswalk. Every entity in our system has a BRR-E- internal ID that links across FFIEC, FDIC, Federal Reserve NIC, NCUA, and SEC EDGAR identifiers.

The timeliness problem. Call report data drops on a lag, often 60-90 days after quarter close. But amendments get filed continuously. Preliminary data gets restated. If you cache a pull from the raw FFIEC files and do not track amendments, you are working with stale numbers. We reload amended filings so that what we serve reflects the most current data the agencies have published.

The result is BankRegReports: clean, validated, cross-source-linked bank regulatory data through a modern API, at a cost structure that works for independent analysts, fintechs, and smaller institutional teams, not just the largest asset managers in the world.


Histogram of return on assets across every FDIC-insured U.S. bank
One API call reaches the whole industry. Return on assets across every FDIC-insured bank is the kind of full-universe distribution the API is built to return. BankRegReports analysis.

What the API Covers

The BankRegReports Data API exposes 101 datasets across 16 domains. All data is sourced from public regulatory filings (FFIEC 031/041/051 call reports, FR Y-9C, UBPR, FDIC Summary of Deposits, NCUA 5300, Federal Reserve NIC, and SEC EDGAR), cleaned, validated, cross-source linked, and served with consistent field naming and units.

Domain Datasets Key sources
Bank directory, snapshot & trends 6 FFIEC call reports, UBPR
Bank sub-books 6 Schedule RC-C, RC-B, RC-E, HMDA
Bank structure & events 5 NIC, FDIC history, FR Y-15
Bank scores & peers 6 Derived + UBPR peer groups
Branches 5 FDIC Summary of Deposits + NIC
Bank risk & compliance 4 ML predictions, enforcement actions, CFPB
SEC / EDGAR 15 10-K, 10-Q, 8-K, XBRL, Form 4, 13D/G, 13F, proxy
Screeners 5 Full-universe + M&A + loan concentration + growth
Industry & macro 10 FFIEC aggregates, Fed data
Rates & yield curve 7 Freddie Mac, Treasury, FDIC national rates
Risk models 4 Failure-prediction model outputs
Failures, events & enforcement 5 FDIC failures, NIC events, enforcement press releases
Credit unions 7 NCUA 5300
Holding companies 3 FR Y-9C
Custom peer groups, watchlist & saved graphs 6 User-configured
Reference & catalog 7 MDRM definitions, metric catalog

Getting Started

Authentication

Get a free API key at bankregreports.com/api. All keys begin with brr_. The free tier covers read access to the core bank snapshot, screener, and industry aggregate endpoints. Paid tiers open the full dataset catalog, including ML predictions and SEC data.

pip install bankregreports
import bankregreports

# Pass the token directly or set BANKREG_API_TOKEN env var
brr = bankregreports.BankReg("brr_xxx")  # or set BANKREG_API_TOKEN env var

# Works as a context manager too
with bankregreports.BankReg("brr_xxx") as brr:
    df = brr.screener(state="GA")

The SDK targets Python 3.8+. It has a hard dependency on requests>=2.25 and an optional dependency on pandas for DataFrame output. Methods return DataFrames by default, which needs pandas; pass as_dataframe=False to any method to get the raw JSON records instead.


The Python SDK: Method Reference

Bank snapshot and trends

# Latest snapshot - all core metrics for one institution
df = brr.bank(852218)
# One row: identity, size, capital, asset quality, and earnings metrics

# 20-quarter time series
df = brr.bank_trends(852218, quarters=20)
# Returns a row per quarter: report_date, all core metrics

# Full Call Report / UBPR detail - ~150 fields
df = brr.bank_deep_dive(852218)
# Includes MDRM-mapped sub-schedules: RC-C loan book, RC-B securities,
# RI-B charge-off and recovery detail, RI interest income waterfall

# Combined profile: snapshot + scorecard + percentile ranks + ML prediction
df = brr.bank_profile(852218)
# Useful when you want the full picture in one call

# Side-by-side comparison of multiple institutions
df = brr.bank_compare([852218, 480228, 37, 92])
# Returns one row per institution, same column schema

The institution identifier used throughout the API is the Federal Reserve RSSD ID, the stable cross-agency identifier that links FFIEC, FDIC, and NIC records. If you have an FDIC certificate number instead, the screener rows carry both rssd_id and fdic_cert, so you can map one to the other.

Screener

The screener is the full-universe filter. The universe covers about 4,300 FDIC-insured commercial banks and savings institutions.

# All banks in Georgia with assets above $500M (assets are in thousands)
df = brr.screener(state="GA", min_assets=500_000)

# Profitable community banks with strong capital
df = brr.screener(
    max_assets=10_000_000,   # under $10B
    min_cet1=12.0,           # CET1 ratio, percent
    min_roa=1.0,             # ROA, percent
    sort="roa",
)

# Well-capitalized banks with low problem loans and a Texas ratio under 10%
df = brr.screener(well_capitalized_only=True, max_npl=0.5, max_texas=10)

The screener accepts state, tier, min_assets, max_assets, min_cet1, max_npl, min_roa, max_efficiency, max_texas, min_nim, well_capitalized_only, and sort. It ignores any other filter name, so check spelling if a filter seems to have no effect. Each row carries rssd_id, fdic_cert, legal_name, hq_city, hq_state, asset_tier, total_assets (thousands of dollars), cet1_ratio, npl_ratio, texas_ratio, roa, roe, nim, efficiency_ratio, leverage_ratio, and is_well_capitalized. CRE concentration screening lives in brr.loan_concentration(). Results are paginated, 50 rows per call by default; see the Pagination section below.

Scorecard and peer comparison

The scorecard produces a CAMELS-style composite grade from five component scores: capital, asset quality, earnings, liquidity, and market risk sensitivity. Each component is scored on a 1–100 scale relative to the institution’s peer group. The composite maps to a letter grade.

df = brr.scorecard(852218)
# Component scores and a letter grade
df = brr.scorecard(852218, composite=True)
# Adds peer-comparison grades

df = brr.peer_comparison(852218)
# The bank against its peer-group benchmarks, about 30 fields in 7 sections

Peer groups are determined by asset tier and charter class, following UBPR methodology. A $2B community bank is compared against other $1B–$10B state-chartered institutions, not against JPMorgan.

Sub-books: deposits, securities, loans, HMDA

# Deposit composition: DDA, MMDA, time deposits, brokered, uninsured
df = brr.bank_deposits(852218, trends=True)

# Securities portfolio: HTM vs. AFS, unrealized gains/losses, OTTI
df = brr.bank_securities(852218)
# Includes AOCI impact on tangible common equity

# Loan portfolio: CRE, C&I, consumer, residential, agricultural concentrations
df = brr.loan_portfolio(852218)
# CRE concentration is reported as % of risk-based capital (regulatory definition)
# Construction concentration is broken out separately

# HMDA mortgage lending data - application counts, approval rates, LTV
df = brr.bank_hmda(852218)

The loan portfolio endpoint maps to Schedule RC-C Part I and Part II. CRE concentration is computed using the regulatory definition from the December 2006 joint guidance: construction and land development loans, multifamily loans, nonfarm nonresidential loans not secured by owner-occupied properties, and loans to finance CRE not secured by real estate, divided by total risk-based capital. This is the number examiners look at, not total CRE divided by assets.

Bank structure and corporate history

# Organizational structure: the parent chain above the bank
df = brr.bank_structure(852218)

# Full structure, including all subsidiaries and equity stakes
df = brr.bank_structure(852218, full=True)

# NIC structural events: mergers, charter changes, openings, closures
df = brr.bank_events(852218)

# FDIC institutional history change log
df = brr.bank_history(852218)

Risk, ML predictions, and enforcement

# ML one-year failure probability, with per-feature contributions
df = brr.prediction(852218, detail=True)

# Banks scored as likely M&A targets
df = brr.ma_screener(
    state="TX",
    max_assets=2_000_000,   # thousands of dollars; defaults to $10B
    min_roa=1.0,
)
# Also accepts min_cet1, max_npl, and min_core_dep

# Enforcement action history: MRAs, CMPs, consent orders, formal agreements
df = brr.enforcement_actions(852218)
# Covers OCC, FDIC, Federal Reserve, CFPB actions

The failure prediction model is trained on 20+ years of FFIEC call report history including the 2008–2012 failure wave, thrift failures, and more recent cases. It is explicitly not built on UBPR peer-group comparisons alone; it incorporates rate environment features, macro-conditioned stress projections, and concentration risk at the loan-category level. SHAP attribution tells you which factors are driving the score for a specific institution.

SEC / EDGAR

# Filing index: 10-K, 10-Q, 8-K, DEF14A
df = brr.sec_filings(852218)

# XBRL financial facts from 10-K/10-Q filings
df = brr.sec_facts(852218, limit=500)

# Form 4 insider transactions (paginated with limit and offset)
df = brr.sec_insider_txns(852218, limit=100)

# 13F institutional holdings - quarterly positions
df = brr.sec_13f(852218)

# Executive compensation from proxy statements (DEF14A)
df = brr.executive_comp(852218)

# 13D/13G significant ownership disclosures
df = brr.sec_13dg(852218)

The SEC data is linked to the FFIEC universe via the BRR-E- entity crosswalk. If you are looking at a bank holding company that files separately as a public company, both the regulatory (Y-9C) data and the SEC filing data are accessible under the same RSSD ID.

Credit unions (NCUA 5300)

# NCUA 5300 trend series for one credit union
df = brr.credit_union(5536, trends=True)
# Uses the NCUA charter number as the identifier

# Credit union directory, filtered by state
df = brr.credit_unions(state="NC")

# CU peer benchmarks by asset tier
df = brr.cu_peer_benchmarks()

Credit union financials follow NCUA 5300 conventions, which differ from FFIEC call report conventions in meaningful ways. Deposits are “shares.” Capital is “net worth.” Loan quality uses “delinquency” not “nonperforming.” The units and terminology in our CU endpoints match NCUA convention, not bank convention.

Holding companies (FR Y-9C)

# BHC snapshot from the FR Y-9C
df = brr.holding_company(1120754)  # uses RSSD ID of the holding company
# Consolidated BHC-level metrics

# Quarterly BHC trend series
df = brr.holding_company(1120754, trends=True, quarters=12)

The Y-9C data covers BHCs with total consolidated assets of $3B or more (and smaller BHCs that opt in). For smaller community bank holding companies that do not file the Y-9C, use the bank-level call report endpoints instead.

Rates, yield curve, and macro

# Current yield curve snapshot
df = brr.yield_curve(latest=True)

# Daily history for one tenor (the 10-year Treasury)
df = brr.yield_curve(tenor="Y10")

# Freddie Mac Primary Mortgage Market Survey
df = brr.rates("mortgage")

# FDIC national deposit rate benchmarks
df = brr.rates("national")

# Treasury bill rates
df = brr.rates("treasury_bill")

# Industry aggregate trend - 20-quarter window
df = brr.industry(trends=True, quarters=20)

rates() accepts mortgage, consumer, deposit, treasury_bill, real_yield, and national.

Watchlist and alerts

# Add an institution to your watchlist
brr.watchlist_add(852218)

# Set a threshold alert - notify when CET1 drops below 10%
brr.alert_create(852218, metric="cet1_ratio", direction="below", threshold=10.0)

# List your alerts
df = brr.alert_list()

Reference and catalog

# Full dataset catalog - 101 datasets: name, path, description, tier
df = brr.list_datasets()

# MDRM code lookup - FFIEC item definitions
df = brr.reference("RCON2170")

# Index of the other reference tables (states, regulators, and more)
df = brr.reference()

# Metric catalog - every available metric with its units and definition
df = brr.metrics()

# Generic accessor - forward-compatible for new datasets
df = brr.dataset("screener", state="GA", min_assets=500_000)

The metric catalog is worth calling out specifically. It lists every metric the API serves with its unit and definition, so you know what a number means before you use it. That kind of methodology transparency is what we think the industry has been missing.


Pagination

List endpoints page with limit and offset query parameters. The SDK wraps this in a generator, brr.pages(), which steps the offset for you and stops at the last page:

import pandas as pd

# Iterate pages manually
for page_df in brr.pages("screener", page_size=200, state="CA"):
    process(page_df)

# Collect all into one DataFrame
all_banks = pd.concat(brr.pages("screener", page_size=500))
print(f"Total: {len(all_banks)} institutions")

Each response’s meta block carries count (rows in this page), total (rows matching the filter), limit, offset, has_more, and next_offset. The DataFrame .attrs["meta"] dict carries this information through. Deep pagination is a paid-plan feature: free and anonymous callers get at most 50 rows per page and cannot page past offset 1,000.


Error Handling

All exceptions subclass BankRegError. GET requests retry automatically on 429 and 5xx with exponential backoff (3 retries by default).

from bankregreports import (
    BankReg, BankRegError,
    AuthenticationError,    # 401 - bad or expired API key
    UpgradeRequiredError,   # 403 - dataset requires higher tier
    NotFoundError,          # 404 - institution not found
    ValidationError,        # 422 - bad parameters
    RateLimitError,         # 429 - throttled; check e.retry_after
    ServerError,            # 5xx - auto-retried
)

try:
    df = brr.prediction(852218, detail=True)
except UpgradeRequiredError:
    print("This dataset requires a paid tier - upgrade at bankregreports.com/api/")
except NotFoundError:
    print("Institution not found - verify the RSSD ID")
except RateLimitError as e:
    import time
    time.sleep(e.retry_after or 60)
    df = brr.prediction(852218, detail=True)
except BankRegError as e:
    print(f"API error: {e}")

Rate limits vary by tier. The free tier allows 200 requests per day. The Pro tier allows 10,000 requests per day. See API pricing for the current plans.


Configuration

brr = bankregreports.BankReg(
    token="brr_xxx",                    # or BANKREG_API_TOKEN env var
    base_url="https://api.bankregreports.com",  # default
    timeout=30,                         # seconds, default 30
    max_retries=3,                      # default 3
)

For local development against a self-hosted instance or test environment, override base_url:

brr = bankregreports.BankReg("brr_dev_xxx", base_url="http://localhost:8000")

Data Quality Philosophy

A few things we do that we think matter and that we have not seen documented consistently elsewhere:

Units are explicit, not assumed. Raw FFIEC call report data mixes thousands, whole dollars, percentages, and basis points across different schedules without always flagging which is which. Every metric in our system has an explicit unit type registered in the metric catalog. When we serve a NIM value of 3.21, that is a percentage. When we serve total_assets of 852218000, that is thousands of dollars (consistent with Schedule RC convention). Nothing is ambiguous.

Calculations are validated against ground truth. We run cross-validation against UBPR peer reports for the full bank universe of about 4,300 institutions across 20 quarters. Any metric that diverges from the UBPR value by more than expected tolerances triggers a formula review. When the UBPR and the raw call report disagree, which happens because the UBPR applies adjustments, we document the divergence in the metric catalog rather than silently choosing one.

Amendments are picked up. Institutions file amendments after the original deadline. We reload amended filings, so the value we serve is the most recent one the agencies have published. The API does not keep an as-originally-filed history, so if your analysis needs first-reported values, archive your own pulls.

Dead codes are filtered. The MDRM has thousands of discontinued codes still present in the bulk files with zero values. We maintain an active/inactive status on every MDRM code and filter out dead-code noise from the data we serve.


Practical Use Cases

A few common patterns we see from API users:

ALCO data package. Pull bank_trends for your institution and peer_comparison for your peer group, join with yield_curve and rates("national"), and you have the rate environment and deposit cost benchmarking data for a standard ALCO rate-sensitivity deck.

Regulatory exam prep. Pull scorecard to identify where your institution’s component scores sit relative to peers. Pull enforcement_actions to build the regulatory history summary. Pull loan_portfolio to get CRE concentration compared with the guidance screening levels (300% of total risk-based capital for CRE, 100% for construction); the API does not apply the guidance’s 36-month growth test, so treat a flag as a prompt for review, not a finding.

M&A screening. Use ma_screener to filter the universe by geography, size, and financial profile. Pull bank_profile on target candidates to get the full picture in one call, including the ML failure-probability estimate.

Credit research. Pull bank_trends for the 20-quarter time series and plot NIM compression, provision expense, and NCO rates. Use bank_securities to quantify unrealized losses on the AFS portfolio. Cross-reference with sec_filings for the 10-K disclosures on interest rate risk.


REST API Directly

If you prefer not to use the Python SDK, the full API is available as a REST API with interactive documentation at api.bankregreports.com/api/v1/docs/. Every endpoint documented above has a corresponding REST path. The Swagger UI lets you explore and test any endpoint directly in the browser with your API key.

Authentication is via an Authorization: Bearer brr_xxx header on every request.

All responses return JSON with a consistent envelope:

{
  "data": [...],
  "meta": {
    "count": 50,
    "total": 4389,
    "limit": 50,
    "offset": 0,
    "has_more": true,
    "next_offset": 50
  }
}

Get Started

Free API key at bankregreports.com/api. The free tier covers enough to build a working prototype: bank snapshots, screener, industry aggregates, yield curve. Upgrade when you need the deeper sub-books, enforcement history, ML predictions, or deep pagination.

If you run into something that does not reconcile with your own data pull from the raw FFIEC files, we want to know. The methodology documentation in the metric catalog is meant to be the definitive explanation of how every number is computed. If it is not, that is a bug.


BankRegReports publishes call report, UBPR, FDIC, Federal Reserve, NCUA, and SEC EDGAR data for every U.S. bank and credit union, with peer benchmarks and quarterly history. Look up a bank