Skip to content

Latest commit

 

History

History
350 lines (272 loc) · 55.4 KB

File metadata and controls

350 lines (272 loc) · 55.4 KB

طبقة · Tabaqa — The Market-Gap Evidence Dossier

Purpose: the "Norges-Bank move" — prove the market and the mechanism with the world's most conservative institutions, so the pitch borrows their credibility instead of asking for belief. Every item below was verified against a fetched primary source (2026-07-04; extended and re-verified 2026-07-13). Several famous numbers were CORRECTED in verification — the landmine list at the bottom is as important as the dossier itself. Use ONLY the numbers on this page. Companions: JUDGE_SCRIPT.md · PROOF.md · DATA_DEFENSE.md

2026-07-13 fold-in. The Jul 8 research run (app/RESEARCH-2026-07-08.md) and the scale brief (SCALE_STORY.md) are now merged here: the funding comparables (Tier 2B), the SAMA legal chain for reading banks and wallets (Tier 5), and the open-source licence map (Tier 7). New this pass: Tier 6 — behavior-priced finance is already live, which is our answer to AMAD's 5th criterion, قابلية التنفيذ الفعلي في القطاع المالي ("can this really ship in a bank?").

2026-07-16 fold-in — the riyals, and one correction that matters. The Jul 15 research run (app/RESEARCH-2026-07-15.md, 104 agents, 22 confirmed / 3 refuted) supplies the SAR figures this dossier deliberately refused to guess at — now sourced to SAMA / FSDP / IMF / GASTAT, in Tier 4A. Two consequences you must absorb before speaking:

  1. The wedge is ~6 million, not ~11 million. Rigorous GASTAT adult (15+) base. Every doc reconciles to 6M — see the reconciliation box under #20. Saying 11M now contradicts our own deck.
  2. "We don't have a market size" is no longer true — landmine #14 is rewritten, not deleted. The SAR 1.4T blog figure stays dead; its primary replacement is SAMA's own SAR 3,186.3bn. New labels [P→C] and [ILL] exist because these figures are computed, not published. Respect them.

HOW TO READ THIS DOC — the honesty labels

Every claim carries its source and its verification status. No exceptions. If it has no label, it does not get said.

Label Meaning How you may speak it
[P] PRIMARY-VERIFIED — fetched from the body that owns the fact (regulator, registry, the paper itself, the company's own press release about itself) State it as fact.
[P→C] COMPUTED-FROM-PRIMARY — arithmetic on primary inputs (added 2026-07-16) State it with the arithmetic on the slide, every input labelled with source + vintage. Never present the output as if a regulator published it.
[S] SECONDARY — reputable trade press, analyst house, or law firm "Industry research puts it at…" — never as gospel, never as regulator data.
[SR] SELF-REPORTED — a company's own marketing / impact number, unaudited "They report…" — or don't say it at all.
[ILL] ILLUSTRATIVE — a scenario resting on an assumption (reduction %, realization %) (added 2026-07-16) Say "illustrative" out loud, in the same breath. Never "measured."
[X] DO NOT SAY — refuted, unverifiable, or licence-locked Silence.

The rule that governs this file: every number is primary-verified, or it does not get said out loud. Its corollary — refuse, don't guess — applies to us on stage exactly as it applies to our adapters.


THE LINE (Norges-Bank-style, for the pitch)

"أكثر المؤسسات تحفّظًا في التمويل عبرت هذا الجسر قبلنا: فاني ماي تعتمد ١٢ شهرًا من التدفق النقدي البنكي لتمويل من لا سجل لهم، والجهات الرقابية الأمريكية الخمس باركت بيانات التدفق النقدي في بيان مشترك، وبنك التسويات الدولية قاسها: 0.76 مقابل 0.64. لا نطلب من السعودية قفزة — نجلب كتاب لعبٍ مُثبتًا إلى السوق الوحيد الذي بنى للتوّ سكّته."

"The most conservative institutions in finance already crossed this bridge: Fannie Mae approves no-score borrowers on 12 months of bank-statement cash flow, all five US banking regulators jointly blessed cash-flow data, and the BIS measured transaction data beating bureau scores 0.76 to 0.64. We're not asking Saudi Arabia to take a leap — we're bringing a proven playbook to the one market that just built the rail for it."

TOP 3 FOR THE 3-MINUTE PITCH

  1. Fannie Mae (Dec 2022) — approves borrowers with no credit score using 12 months of bank-account cash flow; its research: bank-statement cash flow is "more predictive… especially for consumers with no or limited credit history."
  2. BIS 0.76 vs 0.64 — the central bank of central banks measured transaction-data models vs bureau scores; bureau-only lending would have excluded 30% of good borrowers. (Pairs with our own +0.203 ablation — we replicated a BIS-class result.)
  3. The Saudi wedge + the rail78.8% of Saudi adults banked (Findex 2024) vs 56.7% with a credit file (World Bank) — a ~22-point wedge of banked-but-unscorable people — and SAMA built the open-banking rail (2022) and issued its first licences (March 2026). Vision 2030's SME KPI (9.4% → 20%) demands the missing scoring layer.

If the panel leans "can you actually build this in a bank?" → the 4th weapon is Tier 6: Discovery Bank has priced credit on behaviour since 2019, 1.2M customers. Category proven; Kingdom empty.


TIER 1 — Regulators · GSEs · Central Banks

# Source The verified finding Use for
1 FinRegLab — "The Use of Cash-Flow Data in Underwriting Credit: Empirical Research Findings" (Jul 25, 2019). The canonical independent study: 6 non-bank lenders — Accion, Brigit, Kabbage, LendUp, Oportun, and Petal — analysis conducted by Charles River Associates. [P] — fact sheet read verbatim 2026-07-13. Four findings, in the source's own words: ① Predictiveness"The predictiveness of the cash-flow scores and attributes was generally at least as strong as the traditional credit scores and credit bureau attributes studied." ② Combined models — cash-flow data "separate[s] risk in somewhat different ways" and "frequently improved the ability to predict credit risk among borrowers that are scored by traditional systems as presenting similar risk of default"across all traditional score bands. ③ Inclusiveness — among these lenders, 45–50% of borrowers score below ~650. ④ Fair lending"the cash-flow data appeared to provide independent predictive value across all groups rather than acting as proxies for demographic group." Backdrop: 45–60M Americans lack sufficient credit history, while "more than 96 percent of American households have bank or prepaid accounts." Report · Fact sheet (PDF) THE data-layer anchor — the strongest citation in this file. Cash flow predicts at least as well as a bureau score, adds independent signal, and is not a demographic proxy. And read ③+backdrop together: >96% of US households are banked while 45–60M are unscorable — that is the US mirror of our Saudi wedge (#20). The gap we're attacking isn't a Saudi quirk; it's structural, and America proved it first.
23 FinRegLab — "Advancing the Credit Ecosystem: Machine Learning & Cash Flow Data in Consumer Underwriting" (press release Jul 1, 2025) [P] "Overall, the machine learning model that combined cash flow and credit bureau data performed the strongest." At risk cutoffs used by mainstream lenders, the two strongest ML models increased approvals by ~4% over simpler analytics. ⚠️ Disclose the caveat — it is in the source: "data limitations made it difficult to evaluate impacts on consumers who are most likely to benefit because they have little or no traditional credit history." Press release "Cash flow + bureau together beats either alone — FinRegLab, 2025." This is the citation for fusion, which is exactly what Tabaqa does (bank + wallet). ⚠️ Never claim FinRegLab proved the thin-file gain in the ML study — it says it couldn't measure it. Our own +0.203 / +0.117 ablations are where that evidence lives.
2 Fannie Mae DU (Aug 2021 + Dec 2022) [P] 2021: rent from bank statements — 17% of declined applicants would have been approved; <5% of renters have rent on file. 2022: no-score borrowers underwritten on 12-month bank cash flow. 2021 · 2022 THE conservative-institution anchor.
3 Freddie Mac (May+Jun 2022) [P] Underwrites from bank-account data incl. rent paid via Zelle/Venmo/PayPal. AIM · Rent "Both US mortgage giants read wallet payments as credit evidence." — the closest thing to a foreign precedent for our wallet layer.
4 Five US regulators, joint statement (Dec 2019) — Fed+CFPB+FDIC+NCUA+OCC [P] Cash-flow data "may present no greater risks than data traditionally used"; cash-flow evaluation "particularly beneficial" for varied-income consumers. Statement Underused gem: "all five US banking regulators jointly blessed exactly our data."
5 Bank of England (2019) [P] Committed to an "open platform for SME finance" + portable credit file as national infrastructure. BoE UK central-bank endorsement of the concept.
6 CFPB §1033 (Oct 2024) ⚠️ [P] US open-banking rule finalized — but enjoined Oct 2025, being rewritten. Statutory right (2010) stands. Rule Say: "finalized 2024, now being rewritten — direction settled, details not." NEVER "in effect."
7 UK adoption (OBL/FCA) [P] 16.5M users · ~2B API calls/mo · 145 providers (2025); FCA: payments +53% YoY. OBL · FCA "Not a pilot — plumbing."
8 Cambridge CCAF (Nov 2024) [P] 95 jurisdictions with open banking/finance; 54 regulation-led. Report Global inevitability in one number.

TIER 2 — Big-Tech / Market Validation

# Source The verified finding Use for
9 Apple → Credit Kudos (Mar 2022) [P] UK open-banking credit scorer; Companies House shows the company is now literally "Apple Payments Services Limited." (Price ~$150M is press-only — don't state as fact.) Registry "Apple didn't debate the thesis — it bought the company."
10 Visa–Plaid $5.3B · Mastercard–Finicity $825M (2020) [P] Primary releases verified (Visa deal later terminated after DOJ suit — say "bid," not "bought"). Plaid today: 12,000 institutions, 1-in-2 US banked adults. Visa · Mastercard "The giants priced this thesis years ago."
11 Experian Boost (US 2019 · UK Nov 2020) — ⚠️ HANDLE WITH CARE, see the trap below [P/SR] Experian's own study of its own product (Feb 5, 2021): 60% of people who completed Boost saw their FICO Score rise, avg +12 pts; starting below 580 → 87% rose, avg +22 pts; thin-file → 85% rose, avg +19 pts; 47% of previously unscoreable users built enough file to become scoreable. By Jul 2021: 7M+ users, ~50M points added, $1.7B in credit accessed. Study · Scale ONE legitimate use: proof that consumers will consent to share bank data when doing so can only help them. It is NOT a product comparison.
12 UltraFICO / FICO Score XD [P] FICO itself builds cash-flow scores (UltraFICO alive 2026, Plaid-connected; XD: +15M newly scorable). Limited distribution — use as concept proof, not scale. UltraFICO "Even FICO hedged its own bureau moat with cash flow."

⚠️ THE EXPERIAN BOOST TRAP — read this before you cite #11

Boost is consented bank data raising a bureau score. That is precisely the "scoring company" identity Tabaqa pivoted away from on Jul 12. Cite it as a product analogy and you walk straight into:

"So you're Experian Boost for Saudi Arabia?"

— and you have just re-litigated our own pivot, on stage, for free, against ourselves.

The only sanctioned use — consent-willingness, nothing more:

"When sharing data can only help them, people share it: 47% of Experian's unscoreable users became scoreable. Consent isn't the barrier. The barrier is that nobody in the Kingdom turns that consent into a price."

Then move immediately to the difference: Boost hands you a better number and sends it to the bureau. Tabaqa hands you priced offers from competing lenders — and never touches a bureau. Boost feeds the incumbent; Tabaqa routes around it.

Correction logged 2026-07-13: this dossier previously said "~90% of thin-file users gain instantly (avg 19 pts)." The primary study says 85% (avg +19). The 90% was never in the source. Use 85%.

TIER 2B — The money: who funded exactly this thesis

The one-breath line — never a precise category total:

"Venture investors have put $200M+ into exactly this thesis — conservatively; ≈$285M across three companies alone."

Then pivot instantly to Tabaqa. We never pitch Nova, Lean, or anyone else — comparables validate the category in one breath, and every other sentence is about us. No more names unless a judge asks.

# Source The verified finding Use for
24 Nova Credit — $45M Series C (Oct 17, 2023) [P] Canapi Ventures-led. The press release's own headline, verbatim: "…to Scale Cash Flow Underwriting" (the Cash Atlas product). Existing investors: General Catalyst, Index, Kleiner Perkins, Y Combinator. Alive — later raised a $35M Series D. Nuance: use of funds also covered geographic/product expansion, and the C was smaller than the $50M B (CEO framed it as dilution avoidance). PR The cleanest "a VC wrote a cheque for this exact sentence."
25 Lean Technologies (Riyadh) — >$100M total [P] $3.5M seed (Jul 2020) + $33M Series A (Jan 2022 — Sequoia India's Gulf debut) + $67.5M Series B (Nov 11, 2024 — General Catalyst's first-ever KSA investment) = $104M. PR · CNBC "$100M+ of global capital is already building the rail — in Riyadh." ⚠️ Framing discipline: Lean is open-banking infrastructure (A2A payments + data APIs), not a scoring company. Say "infrastructure enabling cash-flow underwriting." Lean is our supplier, not our competitor — and saying so out loud is how we show we know the market.
26 Petal — the mechanism scaled; the card business failed [P] ~400,000 consumers approved since 2018, the majority thin- or no-file at approval, via cash-flow underwriting on customer-permissioned open-banking data (acquirer's PR, verbatim — ⚠️ the 400k is [SR]). $140M Series D Jan 2022 (~$800M valuation) → distressed acquisition by Empower, announced Apr 9, 2024, at a small fraction of that; card retired; acquirer rebranded to Tilt (2025). Prism Data (the 2021 scoring spinoff) survives, independent. Plaid case study: 30% lower roll rates vs bureau-only. PR DISCLOSE PROACTIVELY — see the lesson below.

🎯 THE PETAL LESSON — say it before a judge finds it

Petal is one of the six lenders inside the FinRegLab study (#1). The single best-funded US player proved the mechanism — 30% lower roll rates than bureau-only, ~400k thin-file approvals — and still died as a card issuer: it lent from its own balance sheet, with no wallet fusion and no applicant-facing score.

The line, volunteered — never extracted:

"The mechanism survived; the balance sheet didn't. So: Tabaqa never lends — we price, and lenders fund. We fuse the wallet. And the number belongs to the applicant. Three lessons, three design decisions."

This is the strongest available answer to "what if you're wrong?" — and it is only strong if we raise it first.

TIER 3 — Academic / Central-Bank Research

# Source The verified finding
13 BIS WP 779 (2019) [P] Mercado Libre: transaction-data ML AUROC 0.76 vs 0.64 bureau-only; bureau-only lending would exclude 30% of served borrowers. ⚠️ The famous "0.81 vs 0.71" is NOT in the paper. BIS
14 BIS WP 881 — "Data vs collateral" (2020) [P] 2M+ Chinese firms: big-tech credit reacts to transaction volumes, not collateral — data replaces collateral. BIS
15 Berg et al., RFS 2020 [P] Digital footprint AUC 69.6% vs bureau 68.3%; combined 73.6%. Paper
16 IMF WP 2020/193 [P] 1.8M MYbank loans: fintech model AUC 0.84 vs 0.74 bank scorecard; 0.83 with ZERO credit history — transaction data replaces the missing file. IMF
17 Philadelphia Fed WP 17-17 (2017) [P] LendingClub grade↔FICO correlation fell 80%→35% as alternative data took over; same-risk subprime borrowers got cheaper credit. Fed
18 Norges Bank transaction-data credit study [X] DOES NOT EXIST (5 searches, EN+NO). Never cite it — the viral Norges/NBIM story is about AI productivity, not credit. Use BIS/IMF for the central-bank slot.

TIER 4 — The Saudi Gap, Quantified

# Source The verified finding
19 SAMA Open Banking [P] Framework Nov 2, 2022 (AIS v1, PIS v2); Lab launched Jan 4, 2023; sandbox TSP approvals (Lean, Feb 2025); first open-banking licences issued March 26, 2026. Framework · Licensing
20 The wedge⚠️ CORRECTED 2026-07-16: ~6M, not 11M [P] Account ownership 78.84% (Findex 2024) vs credit-bureau coverage 56.7% (Doing Business 2020, last official figure; registry 0.0%) = 22.14 points of banked-but-unscorable Saudis. [P→C] Applied to the ~27.4M adult (15+) base (GASTAT Population Estimates 2024): ≈ 6.07M people ≈ ~6 million. Q&A caveat: different vintages (2024 banking vs 2020 bureau); DB discontinued 2021. Findex · DB2020 · GASTAT
21 FSDP SME KPI [P] 5.7% (2019) → target 20% by 2030, interim 11% by 2025; actual 9.4% (Q4 2024, FSDP Annual Report on sama.gov.sa). IFC global MSME gap $5.2T (2025 update: $5.7T). Charter
22 CFPB credit invisibles [P] As published (2015): 26M invisible + 19M unscorable ≈ 45M Americans. ⚠️ 2025 technical revision: 13.5M + 29.7M ≈ 43M combined — cite the ~45M combined figure, know the revision. 2015 · 2025

⚠️ THE 6M / 11M RECONCILIATION — settled 2026-07-16, and it is not optional

This dossier said 11 million until Jul 16. It is now ~6 million, everywhere. Both numbers come from the same 22-point wedge — they differ only in the denominator: 11M applied the wedge to the total population; 6M applies it to the ~27.4M adult (15+) base (GASTAT). A credit file is an adult artifact, so 6M is the defensible one, and it is the one a judge can reproduce.

Why this is a Q&A landmine and not a rounding argument: 11M and 6M cannot both be true on stage. If one doc says 11M and the deck says 6M, a judge who spots it doesn't award us the higher number — they stop trusting the other 21. We chose the smaller, harder number. That choice is itself the credibility.

Known stragglers still saying 11M (deck/script assets, outside this dossier's scope — fix before stage): JUDGE_SCRIPT.md, NORTHSTAR.md, TEAM_BRIEF.md, SCALE_STORY.md, KASHF_VS_TABAQA.md, app/web/public/deck.html (+ the iOS copy). EVIDENCE.md and app/SLIDE_NUMBERS.md are reconciled.


TIER 4A — The gap, in riyals (added 2026-07-16 from app/RESEARCH-2026-07-15.md)

This tier reverses a prior refusal. Until Jul 16 this dossier said "we have no primary market size — don't reach for one." That was correct when the only candidate was a blog. It is no longer true: SAMA and the FSDP publish the pool and the target themselves. The rule that replaces the refusal: computed is allowed — but the arithmetic goes ON the slide, every input labelled with source + vintage. A [P→C] number spoken without its math is just a number we made up.

The headline — the safest riyal figure we own

# Figure The verified finding Use for
40 SAR 338bn — the SME financing gap the Kingdom committed to close [P→C] The only "gap" number backed by the government's own published KPI. FSDP sets SME lending at 20% of bank credit by 2030; actual is 9.4% (Q4 2024) — a 10.6-point gap on SAMA's Q2-2025 total credit. The math goes on the slide:
20.0% × SAR 3,186.3bn = SAR 637.3bn (target)
9.4% × SAR 3,186.3bn = SAR 299.5bn (today)
──────────────────────────
gap = 10.6 pts = SAR 337.8bn ≈ SAR 338bn
Inputs: FSDP Annual Report 2024 [P] · SAMA Key Economic Developments Q2 2025 [P]
THE market-gap headline. "A Saudi banking judge cannot dismiss the Kingdom's own KPI." ⚠️ Counter → rebuttal: "You're mixing a Q4-2024 ratio with a Q2-2025 base.""Flagged — but the gap is 10.6 points regardless of vintage; on any 2024–25 base it's SAR 300–340bn, and Tabaqa attacks the thin-file scoring gap directly."

The pool — TAM, all SAMA Q2 2025 unless noted

# Figure SAR Source Label
41 Total Saudi bank credit (the outer pool) 3,186.3 bn SAMA Key Economic Developments Q2 2025 (+15.8% YoY) [P]
42 Consumer / personal finance (directly addressable) 469.8 bn SAMA Q2 2025 (14.7% of credit) — excludes mortgages & cards [P]
43 Banking-sector assets 4,494 bn FSDP Annual Report 2024 (end-2024) [P]
44 MSME lending outstanding 351.7 bn SAMA (2024) [S]
45 Total real-estate loans 922.2 bn SAMA via Arab News (Q1 2025) [S]
46 Auto / vehicle finance (finance-company channel) 25.16 bn SAMA via Arab News (2024, +18.8% YoY, 26% of finance-co credit) [S]

The paired stage sentence (gap + pool in one breath):

"Saudi banks hold SAR 3.19 trillion of credit, including SAR 470 billion of consumer finance — that's the pool Tabaqa plugs into. And the Kingdom has committed to move SME lending from 9.4% to 20%: SAR 338 billion it must unlock, with thin-file underwriting as the bottleneck."

Vehicle-demo sentence (the SAR 150k auto demo sits inside this): "Vehicle finance through Saudi finance companies alone is SAR 25 billion, growing 19% a year — our SAR 150,000 auto decision sits inside exactly this fast-growing, document-heavy flow."

⚠️ SAMA publishes no clean bank-only auto line — SAR 25bn is the non-bank channel only, a conservative floor. Say so if pressed. Upgrade the [S] rows to [P] off SAMA's "Annual Performance of Finance Companies" + Monthly Statistical Bulletin before the deck ships.

The locked-out credit — the wedge, priced

# Figure The verified finding Use for
47 ~6M unscorable → SAR 55–184bn of consumer credit locked out [P→C], assumption-driven — say "illustrative." 22.14 pts × ~27.4M adults (GASTAT 2024) = ~6.07M people. Credit intensity = SAR 469.8bn ÷ 15.53M scorable adults ≈ SAR 30,250/adult. Full parity = 6.07M × 30,250 ≈ SAR 184bn; conservative (30% realization)SAR 55bn. Innovation #1 — the white space, priced. "Saudi Arabia banks 79% of adults but only 57% are visible to a credit bureau — that 22-point gap is roughly 6 million banked, invisible-to-scoring adults, and even at conservative parity it's SAR 55–180 billion of consumer credit they can't access today." ⚠️ Counter → rebuttal: "You're mixing 2024 banking data with 2020 bureau coverage.""Flagged — it's the newest published bureau-coverage figure; the wedge is directionally robust, and Tabaqa scores from consented cash-flow data precisely so bureau lag stops being the gate."

The bank's ROI — avoidable losses

# Figure The verified finding Use for
48 ~SAR 40bn NPL stock → SAR 4–12bn/yr reducible Stock [P→C]: 1.5% NPL ratio (SAMA FSR 2024, for 2023; provision coverage 151%) × SAR 2,584bn loan book (SAMA, end-2023) = SAR 38.8bn ≈ SAR 40bn. Cross-check: 1.5% × IMF-implied SAR 2,670bn (66.7% of GDP, **Table 5**) = SAR 40bn ✓. Reduction [ILL]: cash-flow underwriting cuts roll rates ~30% (Petal/Plaid, #26) → 10% = ~SAR 4bn · 20% = ~SAR 8bn · 30% = ~SAR 12bn. Data #3 / Feasibility #5 — the bank makes money deploying this. "Saudi banks carry roughly SAR 40 billion of non-performing loans; independent evidence shows cash-flow underwriting cuts roll rates ~30% — even a conservative 10–20% improvement is SAR 4–8 billion of avoidable losses a year." ⚠️ Counter → rebuttal: "NPLs are already low at 1.5% — where's the upside?""Exactly why the marginal, thin-file applicant is where losses concentrate — better underwriting reduces NEW defaults among borrowers banks currently reject or misprice, not the existing stock. It's a forward-flow number, labelled illustrative." ⚠️ Vintage-match: 2023 ratio → 2023 loan book. Cite IMF Table 5, never "Table 7." NPL vintage is 2023 — ~2 yrs old for a Jul-2026 pitch; refresh from SAMA FSR 2025 / IMF Art. IV 2025 (CR 25/223) if published.

⚠️ What we still do NOT have — know it before stage. (1) No primary Saudi cost-to-originate, application-abandonment, or time-to-decision figure exists — see #49; do not build a Saudi SAR operational-savings number, it needs two unsourced inputs. (2) SAMA publishes no consumer-finance borrower count and no clean bank-only auto line — average-ticket and bank auto TAM are derived, not primary. (3) Several SAMA/FSDP/IMF PDFs returned 403/404 to automated fetch; figures were confirmed via search-indexed verbatim text + independent corroboration (SPA, Argaam, Zawya, Arab News). Re-open every primary link and confirm the number live before the deck ships.


TIER 5 — The legal chain: reading banks AND wallets, in Saudi, legally

The single question a banking judge is most likely to ask. We answer it with article numbers. Source unless noted: Implementing Regulations of the Payments and Payment Services Law (issued 13/6/2023, in force) — all articles read verbatim on the SAMA rulebook.

# Article / fact The verified finding
27 Art. 6, item 10 [P] "Payment Account Information Services" (AIS) is a standalone licensable activity — alongside payment initiation and e-money issuance.
28 "Payment Account" is provider-agnostic [P] + definitional inference. Any account used to execute payment transactions; "Payment Account Service Provider" expressly includes but is not limited to licensed banks → EMI wallets (urpay, Barq) and digital banks (stc bank, D360) fall inside AIS-readable scope by definitional chain (mirrors the settled PSD2 reading). ⚠️ Label it honestly if pressed: the regs never literally say "e-wallet = Payment Account." No source says bank-only — but this leg is an inference, not a quote. Say so. That admission is cheaper than being caught.
29 Art. 96(1) [P] Account providers MUST grant licensed AIS providers access on customer consent, on an "objective, non-discriminatory and proportionate basis." — i.e. the bank cannot refuse us if the customer says yes.
30 Art. 23 / Art. 14 — our own on-ramp, priced [P] The PAIS licence carries a SAR 20,000 issuance fee (Art. 23) and — unlike major PI / major EMI / micro EMI — no joint-stock-company requirement (Art. 14). ⚠️ The regs are silent on minimum capital: never say "no capital requirement."
31 First AIS licences issued [P] March 26, 2026 — Neotek ("New Technology for Software Solutions") + Lean Technologies Saudi Arabia, per the Saudi Press Agency. Preceded by the SAMA regulatory sandbox, where Lean's AIS products served lending, insurance and marketplace clients (~1M bank accounts verified — [SR], Lean's own figure). SPA
32 Consent layer — SAMA policy + PDPL [P] SAMA's Open Banking Policy, verbatim: analysis of customers' financial-transaction data "with customer consent — and offer tailored products"; open banking "will expand access to credit"; the standard is "explicit and informed consent." Aligned with PDPL Art. 24 (explicit consent for credit data; effective 14 Sep 2023, grace ended 14 Sep 2024). ⚠️ REQUIRED SOFTENING: this is aspirational policy language. Say "explicit SAMA policy endorsement"NEVER "regulatory blessing." Policy PDF
33 Tarabut — the sandbox-to-certification datapoint [P/S] Sandbox test permit Nov 2022 (single-sourced to Tarabut's own PR) → KSA Open Banking certification, May 30, 2023 (multi-sourced), launching AIS inside the sandbox. ⚠️ Precision traps: this was a framework-compliance certification + sandbox operation — NOT a full licence (those came Mar 2026); and "among the first" holds only as first open-banking fintechs — the SAMA sandbox has run since 2018 (45 permits by Q1 2023). Tarabut

The scale ladder — priced, dated, quotable

Every rung below is a row in the table above. This is the answer to "how do you scale?" — and to "is this actually legal?" — at the same time.

Phase Channel Legal basis Status
Today Customer-uploaded statements (bank + wallet exports) PDPL explicit consent — no licence needed In production (tabaqa.vercel.app)
Growth Plug into a licensed AIS aggregator as a client Art. 96 — providers must grant licensed access on consent Rails licensed Mar 2026; our ingestion is source-agnostic → a new rail is an adapter profile, not a rebuild
Scale Our own AIS (PAIS) licence Art. 23 — SAR 20,000, no joint-stock requirement Priced, dated, quotable

🎤 The judge-proof spoken answer (every beat maps to a verified row above)

"Three consented channels connect banks and wallets. First, a SAMA-licensed AIS aggregator — the exact activity SAMA licensed for the first time on March 26, 2026 (Neotek and Lean, after Lean verified ~1M accounts in the sandbox, including for lenders). Second, customer-permissioned statement upload, which our adapters parse today — legal under PDPL because the customer hands us their own data with explicit consent. Third is screen-scraping, which we refuse and the licensing regime is killing. Is it legal for wallets too? Yes, by statute: the Payments Law defines a Payment Account provider-agnostically — any account executing payment transactions, expressly including but not limited to banks — so urpay and Barq (EMIs) and stc bank and D360 (digital banks) are in AIS scope, and Article 96 obliges providers to grant licensed AIS access on customer consent. Our own path: the dedicated AIS licence costs SAR 20,000, with no joint-stock requirement — or we ride a licensed aggregator until then."

Volunteer if pressed on wallet APIs: the mandatory OB Framework APIs reached banks first — a live wallet-API pull is a rollout question, which is exactly why the statement-upload channel we already shipped matters today.


TIER 6 — "Can this actually ship in a bank?" — behavior-priced finance is already live

Why this tier exists. AMAD's 5th criterion is قابلية التنفيذ الفعلي في القطاع الماليcan this really run in the financial sector? Our answer is not a promise or a roadmap. It is products already in production, in regulated markets, with regulators watching and customers on the books. The category is proven. The Kingdom is the empty part.

# Source The verified finding Use for
34 Discovery Bank (South Africa) — "Vitality Money" [P] Behavior-priced banking, live since March 2019 (a licensed, digital-only retail bank), ~1.2M customers (as of May 2025). Interest rates flex monthly with financial behaviour, measured across six behaviours (planning, savings, short-term debt, insurance, retirement, property). From discovery.co.za, verbatim: "Up to 7% less on your optional single credit facility," up to 5.25% on demand savings, up to 3.5% on everyday balances. Your borrowing rate is a function of how you handle money. EY's case study calls it "the world's first behavio[u]ral bank"; the UN's UNSGSA (Queen Máxima) profiled it (Jun 4, 2025). Discovery · EY · UNSGSA THE feasibility anchor. "A licensed retail bank has priced credit on behaviour for seven years, for 1.2 million customers. This isn't a hypothesis — it's a product with a P&L, a regulator, and a UN case study." ⚠️ Two hard rules before you open your mouth — see the box below.
35 Usage-based / telematics insurance [S] — industry & market-research sources, NOT regulator data. Label them as such. Consented behavioural data → a personalized price, at scale, in a regulated industry: 21M+ US policyholders shared telematics data with their insurer in 2024 (IoT Insurance Observatory, via insurance trade press; ~28% CAGR since 2018). 14.4% of personal-lines motor policies are telematics — globally (Research & Markets / GlobeNewswire, Jun 13, 2025, from a 2024 consumer survey). ⚠️ On discounts, tell the truth: carriers advertise 10–30%+, but the Maryland Insurance Administration (a state insurance regulator) found only 31% of enrolled drivers actually saw their premium go down in 2023; the Consumer Federation of America calls the advertised savings "a mirage for many drivers, or… highly exaggerated." GlobeNewswire · CFA The adoption is the point — not the discount. "Consented behavioural pricing is already normal: 21 million Americans hand their insurer their driving data for a better price. We're doing the same for credit — except the applicant holds the consent, and the price comes back as an offer." ⚠️ If a judge pushes on the discount, concede it immediately"advertised, not realized; a US state regulator found only 31% of drivers actually saved." The concession wins more than the number would have. It also sets up our own rule: we show a real price, not a teaser rate.
49 Freddie Mac — 2024 Cost to Originate Study (added 2026-07-16) [P] — but US single-family mortgage, NOT Saudi. Label it every time. Verbatim: personnel expenses are 67% of total loan-production cost; underwriting "remains fairly manual and therefore costly"; high-automation lenders originate loans "$1,500, or 14%, less costly"; executives "believe it can help them save up to 40%." Study (PDF) The efficiency supplement — never the load-bearing feasibility claim. "Two-thirds of loan-origination cost is people doing manual document and underwriting work; lenders who automate cut cost-per-loan 14% measured, up to 40% believed — Tabaqa delivers that straight-through, in seconds." ⚠️ Quote the measured 14%, not the 40% aspiration. ⚠️ Counter → rebuttal: "That's US mortgage data, not Saudi auto finance.""Correct — I present it as a global efficiency benchmark, not a Saudi figure; the mechanism (manual underwriting + document verification IS the cost) is channel-agnostic, and I quote the measured 14%." 🚫 Do NOT fabricate a Saudi SAR operational-savings number — no Saudi cost-to-originate exists. Feasibility rides on #34 Discovery (live 7 yrs, 1.2M) + Tier 5's legal chain; this row is the supplement.
(11) Experian Boost — cross-reference only [P/SR] 47% of previously unscoreable users became scoreable once they consented to share bank data. Consent-willingness evidence ONLY. Never a product analogy — see the Experian Boost trap at Tier 2.

⚠️ NEVER SAY "FIRST IN THE WORLD"

Discovery Bank exists. Behavioral pricing of credit is a proven category with a seven-year-old licensed bank and 1.2M customers sitting in it. Claim "world-first" and any judge who knows Discovery — or who has read the EY case study, or the UN's — has caught us inflating. And once we're caught inflating one number, every other number we said gets re-examined. That is how a winning deck loses a room.

The correct framing — and it is the stronger one anyway:

"The category is proven — Discovery Bank has priced credit on behaviour since 2019. What does not exist is a Saudi version: on Saudi rails, reading a Saudi wallet, in Arabic, handing the applicant competing offers. Category proven; Kingdom empty."

AR: «الفئة مُثبتة — بنك Discovery يُسعّر الائتمان على السلوك منذ ٢٠١٩. غير الموجود هو نسخة سعودية: على قنوات سعودية، تقرأ محفظة سعودية، بالعربية، وتُعطي المتقدّم عروضًا متنافسة. الفئة مُثبتة، والسوق السعودي فارغ.»

Why this reframe actively raises our feasibility score: "nobody has ever done this" reads to a banking judge as RISK. "This works in three markets; nobody has built it here" reads as EXECUTION. We want to be the second sentence.


TIER 7 — Open source: what we may cite, what we may build on, what we must never publish

# Repo / dataset Verified status Verdict
36 OBP-API (OpenBankProject) [P] THE citable open-source open-banking / AIS server (Open Banking / XS2A / PSD2 / Open Finance; ships sandbox data import; real bank sandboxes built on it — Danske, BNP Paribas, OP). Active (~1.7k stars, Scala). Dual-licensed AGPL v3 / commercial (TESOBE GmbH) — hosting a modified instance triggers AGPL network copyleft. Repo CITE as prior art. NEVER embed.
37 toad (amphibian-dev) [P] MIT — licence verified via three independent primary artifacts (GitHub API SPDX field, LICENSE file, setup.py). Python credit-scorecard toolkit, actively maintained. Repo Safe to build on. Complements our scorecardpy / OptBinning stack.
38 bankstatementparser [P] Apache-2.0. Parses CAMT.053 / ISO 20022, PAIN.001, CSV, OFX/QFX, MT940 + digital and scanned PDFs into a unified Transaction model — the closest open-source comparable to our adapter layer. README grep for arabic|rtl|عرب returned ZERO hits. Repo The Arabic/RTL/Hijri gap is UNCONTESTED — and now verified, not assumed. We may say "we checked." ⚠️ Don't oversell it as a mature rival (33 stars, v0.0.x), and be honest that its LLM/vision PDF path could incidentally read an Arabic PDF — but nothing in it targets Arabic headers, Hijri dates, or RTL.
39 Kaggle — Home Credit Default Risk [P] Rules read verbatim: a competition-specific override supersedes General Rule 7.A ("only for the purposes of the Competition" — stripping the usual academic carve-out), and Rule 7.B separately bans transmitting / duplicating / publishing / redistributing the data. Third-party GitHub and paper usage reflects non-compliance, not permission. Rules 🚫 DOUBLY LOCKED — NEVER in the deck, the demo, the repo, or a sentence. Our external validity comes from Berka + UCI + AlfaBattle instead.

⚠️ THE LANDMINE LIST — corrections that save you in Q&A

A · Numbers that circulate widely and are WRONG (we verified the primary sources)

  1. BIS is 0.76 vs 0.64 — not "0.81 vs 0.71."
  2. Saudi account ownership is 78.8% (Findex 2024) — not "90%+."
  3. CFPB 1033 is NOT in effect — finalized Oct 2024, enjoined Oct 2025, being rewritten.
  4. No Norges Bank credit study exists — the viral story is NBIM + AI productivity.
  5. Experian Boost UK launched Nov 2020 — not 2022.
  6. Lean's SAMA licence is March 2026 — not Jan 2024.
  7. Apple/Credit Kudos price is press-reported — the registry rename to "Apple Payments Services Limited" is the bulletproof fact.
  8. Visa "bid" $5.3B for Plaid — deal terminated after DOJ suit; never say "bought."
  9. Experian Boost thin-file gain is 85% (avg +19 pts) — not "~90%." (Our own error, caught and corrected 2026-07-13. The 90% was never in Experian's study.)
  10. Telematics: "~20% of US auto policies" is NOT VERIFIED — do not say it. ⚫ The 14.4% figure is global personal-lines motor policies, and the separate 20.9% is global customers with a pay-as-you-go policy — a different metric, and not US. Conflating them is how the "~20% of US policies" line was born. What IS supportable [S]: "21M+ US policyholders shared telematics data with their insurer in 2024." Use that, and label it industry research.
  11. Telematics discounts are ADVERTISED, not realized. "10–30% off" is a maximum, not a typical outcome — the Maryland Insurance Administration found only 31% of enrolled drivers saw any decrease (2023). Say "advertised up to 30%; a state regulator found most drivers didn't actually save" — and let the honesty do the work.

B · 🔴 Numbers that are REAL but NOT OURS TO QUOTE (self-reported / unaudited)

  1. 🚫 Discovery Bank's "Diamond-status customers are 97% less likely to be in arrears" — DO NOT QUOTE. It is Discovery's own marketing number, unaudited. It appears in the UN/UNSGSA feature — but the UN is repeating Discovery, not verifying it: every statistic there is attributed to Discovery Bank itself, with no independent source. Repeating it unverified breaks our own evidence rule — the one rule this entire document exists to enforce. DO NOT QUOTE unless primary-verified. Bonus reason to leave it alone: it is almost certainly a selection effect, not a treatment effect. Customers who reach Diamond status are different people from those who don't — the number tells you who they already were, not what the product did to them. A sharp banking judge will say exactly that, and if we're the ones who put the number on the slide, we lose the whole Data axis in one sentence. Cite Discovery for the mechanism being live (#34), never for its size.
  2. Petal's ~400k approvals and Lean's ~1M verified accounts are company PR figures — consistent across independent coverage, but unaudited. Say "they report…".
  3. ⚫ "KSA retail + personal finance ≈ SAR 1.4 trillion" — STAYS DEAD. Do not present this as a fact. Its trail is Ken Research → a WordPress blog — i.e. secondary at best, and the chain of custody is a blog. It has no place in a dossier whose entire claim to authority is primary verification. It survives, correctly labelled as an estimate, in app/PRICING_ENGINE.mdit does not survive here, and it never goes on a slide or into a sentence spoken to a judge. ⚠️ UPDATED 2026-07-16 — the refusal is retired; the replacement is primary. We no longer dodge market-size questions, because SAMA publishes the number itself: SAR 3,186.3bn total bank credit (Q2 2025), of which SAR 469.8bn consumer finance — and SAR 2,584bn (2023) for vintage-matched work (Tier 4A, #41–42). Answer with SAMA's figure, never the trillion. The redirect, now upgraded: "I won't quote a blog's trillion. SAMA's own number is SAR 3.19 trillion of bank credit, SAR 470 billion of it consumer finance. And the gap is the Kingdom's own KPI: SME lending 9.4% → 20%SAR 338 billion it has committed to unlock. Meanwhile 78.8% of Saudi adults are banked and only 56.7% have a credit file — about 6 million people who cannot be priced today."

C · 🗣️ Words that get us caught

  1. NEVER "first in the world" / "world's first" about behavioral pricing — Discovery Bank exists (see Tier 6). Say "category proven, Kingdom empty."
  2. NEVER "regulatory blessing" for data analysis — SAMA's policy language is aspirational. Say "explicit SAMA policy endorsement."
  3. NEVER "no capital requirement" for the PAIS licence — the regs are silent on capital. Say "SAR 20,000 issuance fee, no joint-stock requirement."
  4. NEVER call stc bank / D360 "wallets" or "EMIs" — they are licensed digital banks. The EMI wallet examples are urpay and Barq.
  5. NEVER cite the Jan-2021 SAMA OB Policy for dates or scope — go-live slipped and it contains zero AIS/wallet/EMI content (full-PDF grep). Cite the Nov 2022 Framework + the 2023 Implementing Regulations.
  6. NEVER say Tarabut was "licensed" or name its certification category ("AIS Retail Provider" is unconfirmed). Say "KSA Open Banking certification (May 2023), operating AIS in the sandbox."
  7. NEVER state a precise category funding total. Say "$200M+ into exactly this thesis — conservatively."
  8. NEVER cite TomoCredit as a success comparable — the bureaus cut off its reporting access in Oct 2024; unpaid-bill lawsuits; a Feb 2026 Forbes exposé of its credit-boosting service. Omit it, or use it only as a cautionary tale.
  9. NEVER claim FinRegLab's 2025 ML study proved the thin-file gain — the report says the opposite: "data limitations made it difficult to evaluate impacts on consumers who are most likely to benefit." Our own ablations are where that evidence lives.
  10. NEVER claim bankstatementparser "uses Plaid's 13-category schema" — REFUTED. Don't use that repo to argue the Plaid taxonomy is a de-facto standard.
  11. NEVER cross-quote our own lift numbers. Berka +0.203 = statement-like features; AlfaBattle +0.117 = the mechanism at scale on a card stream. They are different experiments — quote each in its own context, never merged into one number.

D · 🚫 Data we must never publish

  1. Kaggle Home Credit Default Risk — doubly licence-locked (see #39). Never in the deck, the demo, the repo, or a sentence.

E · 💰 The riyal landmines (added 2026-07-16 — killed in the Jul 15 verification pass)

  1. ⚫ "SAR 932.8bn / 29.3% real estate, Q2-2025" — REFUTED (1-2 in verification). Use the confirmed SAR 922.2bn (Q1-2025) instead (#45).
  2. ⚫ "MSME SAR 258bn end-2023" and the derived "SME gap SAR 259bn" — REFUTED (0-3, unanimous). Use MSME SAR 351.7bn (2024) (#44) and the 10.6pp → SAR 338bn gap (#40) instead.
  3. ⚫ "IMF: household DSTI ~40% / data scarce" — REFUTED (0-3, unanimous). Do not cite it at all.
  4. ⚫ "~11 million unscorable Saudis" — SUPERSEDED BY OUR OWN CORRECTION. Say ~6 million. This was our number until Jul 16; it applied the 22-point wedge to the total population instead of the ~27.4M adult (15+) base. The rigorous figure is ~6.07M ≈ 6 million (#20, #47). Every doc reconciles to 6M. Saying 11M now doesn't just overstate — it contradicts our own deck, and the contradiction is what a judge punishes.
  5. ⚫ NEVER present a Saudi SAR operational-savings number. No primary Saudi cost-to-originate exists; building one needs two unsourced inputs (loans/yr, cost/loan). The Freddie Mac 14% (#49) is a US mortgage benchmark — say so in the same breath, and quote the measured 14%, never the 40% aspiration.
  6. ⚫ NEVER say "measured" about the SAR 4–12bn reducible losses. The stock (~SAR 40bn) is [P→C]; the reduction % is [ILL] — an assumption. Say "illustrative."
  7. ⚫ NEVER speak a [P→C] number without its arithmetic. SAR 338bn, ~6M, SAR 55–184bn and SAR 40bn are computed, not published. The math goes on the slide, inputs labelled with source + vintage. A computed number spoken bare is indistinguishable from one we invented — and a judge will treat it as exactly that.

⚠️ STANDING CAVEATS — the things only time can fix (disclose them; don't be caught by them)

  • Wallet-in-scope is a definitional-chain inference. The Payments Law defines "Payment Account" provider-agnostically and Art. 96 imposes the access duty — but no source literally says "e-wallet = Payment Account." No source says bank-only either. If pressed: say it's an inference, and say why it's the settled PSD2 reading.
  • Live wallet-API pull is a rollout question. The mandatory OB Framework APIs reached banks first. This is precisely why our statement-upload channel exists and is in production today.
  • Self-reported figures (Petal ~400k, Lean ~1M, every Discovery statistic) are unaudited. Prefix with "they report."
  • Cold-start honesty line — rehearse it, and volunteer it:

    "Our score is mechanism-validated on a million foreign outcomes and not yet Saudi-calibrated. We know that. The pilot's retro-validation closes it in 60 days."

Created 2026-07-04 from a 5-agent adversarially-verified research pass. Extended 2026-07-13: folded in app/RESEARCH-2026-07-08.md (109-agent run, 23 primary-verified claims) + SCALE_STORY.md; added Tier 6 (behavior-priced finance, live) and Tier 7 (open-source licence map); corrected the Experian thin-file figure; demoted the SAR 1.4T market size. Primary URLs inline. Extended 2026-07-16: folded in app/RESEARCH-2026-07-15.md (104-agent run, 22 confirmed / 3 refuted). Added Tier 4A — the gap in riyals (#40–#48: SAR 338bn SME gap, SAR 3.19tn / 470bn TAM, ~6M → SAR 55–184bn, ~SAR 40bn NPL → SAR 4–12bn) + #49 Freddie Mac cost-to-originate (67% personnel, 14% measured). Added labels [P→C] and [ILL]. Added landmine section E (#27–#33). Corrected the wedge from ~11M to ~6M (GASTAT adult base) and retired the "we can't size the market" refusal — SAMA publishes it. Primary URLs inline; re-open each before the deck ships.