Unmasking Undiagnosed Rare Disease Cohorts
Applying Positive-Unlabeled (PU) machine learning to fragmented claims datasets to identify high-probability patient cohorts where true negative labels do not exist.
Client: Mid-Cap Rare Disease InnovatorStrategic analogs detailing how we isolate commercial bottlenecks, apply deterministic analytical methodologies, and mandate executive action across the life sciences ecosystem.
Applying Positive-Unlabeled (PU) machine learning to fragmented claims datasets to identify high-probability patient cohorts where true negative labels do not exist.
Client: Mid-Cap Rare Disease InnovatorUtilizing sequence-based machine learning on longitudinal Rx/Dx telemetry to predict disease progression 6-12 months before standard clinical diagnosis.
Client: Top 10 Global Neurology LeaderDeconvoluting regional patient hand-offs and applying graph analytics to locate hidden community prescriber influence sinks driving biologic uptake.
Client: Global Tier-1 PharmaIntegrating sub-national P&T formulary approvals and HCP trial-decay logic to predict launch stall curves, enabling decisive resource reallocation at Week 12.
Client: Top 5 Immunology InnovatorMoving beyond naive deciling by estimating heterogeneous treatment effects to direct expensive rep visits exclusively to persuadable targets.
Client: Cardio-Metabolic FranchiseConstructing a deterministic journey waterfall to pinpoint the exact days where diagnosed patients abandon treatment due to administrative friction.
Client: Precision Oncology BiotechDeploying stochastic market simulation to evaluate Phase II assets against future standard-of-care shifts, averting a $400M sunk-cost commitment.
Client: Mid-Cap Autoimmune DeveloperImplementing contract-level profit mapping to renegotiate major PBM tiers, improving net margins by 12% in a high-rebate environment.
Client: Global Pharma EnterpriseUtilizing process mining on site certification logs to cut clinical onboarding times by 50%, allowing the manufacturer to reach peak capacity two years ahead of forecast.
Client: Top 3 CAR-T ManufacturerIn rare disease markets (e.g., lysosomal storage disorders or rare oncologic mutations), official ICD-10 diagnostic coding is chronically delayed or entirely absent in real-world data. Commercial teams struggle to deploy field forces efficiently because they cannot definitively identify where undiagnosed patients reside within regional health systems.
Traditional supervised machine learning requires both confirmed 'positive' and confirmed 'negative' patients to train a predictive model. However, in claims data, an un-coded patient is not definitively 'negative'—they are merely 'unlabeled'. Forcing traditional binary classifiers onto this environment leads to severe model bias, artificially suppressing the identification of legitimate, symptomatic patients.
We deployed an Elkan-Noto based Positive-Unlabeled (PU) Learning framework across a longitudinal claims dataset encompassing over 50 million covered lives. Rather than assuming un-coded patients were negative, the mathematical model treated them as a mixed distribution of hidden positives and true negatives. We isolated highly confident 'true negatives' (patients with mutually exclusive terminal diagnoses) and trained a Gradient Boosting ensemble. This system assigned a latent probability score to the remaining 'unlabeled' mass based purely on longitudinal symptom and minor-procedure phenotypes.
The PU framework successfully identified 4,200 "look-alike" high-risk patients lacking the formal ICD-10 code but displaying identical temporal symptom trajectories to confirmed patients. Commercial leadership instantly routed these unmasked HCP targets to specialized field teams, yielding a measured 31% uplift in formal diagnostic testing requisitions within a 90-day sprint.
For progressive neurological therapies, early intervention is critical to preserving irreversible functional loss. The optimal time to engage a physician with unbranded disease awareness is months before the patient exhibits late-stage, undeniable symptoms.
Physicians frequently fail to officially diagnose the condition until severe symptoms emerge. Commercial marketing efforts tied solely to formal diagnostic claim codes were activating far too late in the patient timeline, resulting in suboptimal clinical outcomes and massive lost market share to established generic stop-gaps.
We bypassed simple cross-sectional, static models and engineered a Recurrent Neural Network (RNN) architecture designed specifically for sparse, irregular time-series healthcare data. By embedding thousands of distinct medical interventions, lab orders, and minor procedural codes as a timeline rather than a snapshot, the model learned the complex temporal sequence signatures that consistently precede clinical onset by 8 to 12 months.
Marketing strategy successfully pivoted from 'post-diagnosis' engagement to 'pre-disease' unbranded awareness. By deploying highly targeted peer-to-peer (P2P) education to HCPs actively managing flagged high-probability cohorts, the brand accelerated diagnosis timeframes by an average of 4.2 months across the targeted population.
A specialized oncology field team was deploying vast, highly expensive commercial resources against high-decile Key Opinion Leaders (KOLs) at major academic medical hubs, operating on historical industry assumptions about influence and prescription volume.
Localized, longitudinal script tracking revealed a structural gap: these academic KOLs rarely wrote the initial prescription. They were acting as surgical or second-opinion validators. Following validation, the actual treatment initiation and ongoing script volume fell back into a highly diffuse, untracked web of regional community hematologists who were largely ignored by the sales force.
We ingested millions of deterministic claims pathways to construct a weighted, directed graph of patient movement across the US healthcare system. Applying a modified PageRank algorithm alongside Louvain community detection, we bypassed basic decile math to isolate the 'true influence sinks'—highly connected community oncologists who routinely received validated patients from KOLs and initiated the long-term infusion regimens.
The commercial blueprint was rewritten entirely. Tier-1 access resources were shifted away from academic validators (retaining only medical affairs liaison support) and directly embedded into the 15% of community nodes that controlled 80% of actual treatment initiation weight, drastically reducing script abandonment post-KOL consult.
At Week 12 of a highly scrutinized immunology launch, top-line NBRx (New-to-Brand Prescriptions) appeared to perfectly match expected street targets. Conventional executive dashboards presented to the board indicated a successful, on-track launch phase.
Standard volume reporting masked a critical underlying dynamic: heavy initial 'trialing' by early-adopter physicians was hiding a near-total failure of 'repeat' prescribing. Without repeat scripts establishing a base, mathematical reality guaranteed a severe trajectory cliff early in Quarter 2, putting the asset's entire fiscal year at risk of missing Wall Street guidance.
We stripped out static linear forecasts and instituted a Bayesian hierarchical model. This model integrated early leading indicators: regional P&T committee meeting velocity, localized access hurdles (extracted from Prior Auth rejection logs), and HCP trial-to-repeat conversion decay rates. The resulting simulation deterministically proved the launch curve would stall at 40% below estimates due to specific regional administrative access friction dampening repeat clinical adoption.
Armed with quantitative proof of the impending stall, leadership executed a "Code Red" field redirection. Dedicated market access personnel were stripped from low-friction zones and flooded into localized P&T friction hotspots, smoothing authorization pathways. The intervention averted the forecasted cliff, securing the asset's $120M annual recurring run-rate trajectory.
Traditional commercial deployment in pharma universally relies on simple volume deciling: identifying the highest historical volume prescribers in a territory and mandating maximum physical sales rep frequency against them.
Volume does not equal persuadability. This legacy model results in massive wasted operational spend by calling on "Sure Things" (brand loyalists who prescribe regardless of rep interaction) and "Lost Causes" (competitor loyalists who will never switch). The executive objective must shift from volume prediction to response elasticity prediction.
We applied a Double Machine Learning causal inference framework to historic CRM interaction logs joined with longitudinal Rx outputs. By isolating and estimating the Heterogeneous Treatment Effect (HTE) of a physical field visit versus a cheaper digital/email interaction, the model assigned an "Uplift Score" to every targeted physician. Non-linear integer programming then mapped the constrained field force strictly to the highly elastic "Persuadable" quadrant.
The revised commercial blueprint isolated 22% of historical Tier-1 targets as "Sure Things" and shifted them entirely to automated digital channels with zero negative impact on volume. The unlocked field capacity was hyper-concentrated onto mid-decile "Persuadables," generating a measured 18% net-new Rx volume lift with no additional commercial headcount.
A specialty oncology brand targeting a rare mutation observed highly acceptable upstream diagnostic testing volumes, but significantly lower-than-expected commercial fulfillment rates. The conversion pipeline from a positive Next-Generation Sequencing (NGS) reflex test to an active, reimbursed commercial patient was failing.
Standard syndicated market data provided no granularity into the complex 6-week onboarding window. Leadership could see patients entering the funnel (Dx) and the few exiting it (Rx), but had zero visibility into exactly where, when, and why the vast majority of patients were abandoning therapy in the middle of the process.
We engineered a deterministic patient journey waterfall utilizing rigorous survival analysis techniques. By linking probabilistic ICD-10 diagnostic coding, lab NGS reflex testing logs, and downstream 867 specialty pharmacy feeds via a tokenized privacy-safe environment, we calculated exact Kaplan-Meier abandonment probabilities at every micro-step of the patient access journey.
The data proved failure was administrative, not clinical. The critical attrition vector occurred almost entirely between Prior Authorization (PA) Appeal 1 and Appeal 2. Leadership executed an immediate deployment of targeted Field Reimbursement Managers (FRMs) to these specific clinic clusters, reducing administrative time-to-therapy by 14 days and salvaging $22M in annual at-risk patient initiations.
A mid-cap biopharma was evaluating whether to progress two distinct autoimmune assets from Phase II to costly, pivotal Phase III trials. The board required a definitive valuation to authorize the $400M clinical expenditure.
Standard Net Present Value (NPV) spreadsheets rely on static assumptions regarding future market share and pricing. However, the autoimmune space is highly volatile, with multiple competitor mechanisms of action expected to read out data before our client's assets reached market. Static math ignored the dynamic reality of a shifting standard-of-care.
We replaced the static models with a stochastic market simulation (Monte Carlo). We mapped the clinical profiles of 15 competing pipeline assets, assigned probability distributions to their trial outcomes based on mechanism validity, and simulated the future market landscape 10,000 times. We tested our client's assets against these thousands of potential future realities to find their true expected commercial value under competitive duress.
The simulation revealed that Asset A had a 70% probability of launching into a market where its efficacy was no longer standard-of-care, effectively rendering its commercial value near zero. Asset B proved highly robust across all scenarios. The board confidently divested Asset A, averting a $400M sunk-cost failure, and doubled clinical investment to accelerate Asset B.
A blockbuster immunology drug was maintaining high prescription volumes, but the executive team noticed severe stagnation in net revenue. To maintain preferred formulary status, commercial teams had negotiated increasingly aggressive rebates with Pharmacy Benefit Managers (PBMs).
The Gross-to-Net (GTN) bubble had become opaque. Because rebates were negotiated in silos across different payer channels (Commercial, Part D, Medicaid), finance leadership could not easily trace the true profitability of a prescription written by a specific physician, filled at a specific pharmacy, under a specific PBM contract.
We built a deterministic, contract-level profit mapping engine. By tying together claim-level dispensing data with exact historical contracting terms and channel-specific rebate obligations, we were able to calculate the true Net margin of every single dispensed script. This highlighted specific 'negative margin' geographies where the company was essentially paying the system to dispense the drug.
The analysis empowered the market access leadership team to walk away from two highly unprofitable regional contracts and renegotiate three major national PBM tiers based on precise elasticity thresholds. While total prescription volume dropped by 3%, net revenue margins improved by 12% across the franchise.
A leading manufacturer was launching a curative CAR-T therapy. Because of the complex manufacturing and handling requirements of autologous cell therapies, hospitals must undergo a rigorous, multi-month certification process before they are legally allowed to prescribe and administer the drug to patients.
The clinical site onboarding process was stalling heavily, averaging over 90 days per hospital. Because patient capacity is strictly limited by the number of certified centers, this operational bottleneck directly choked the launch trajectory, pushing peak sales forecasts back by years.
We deployed digital Process Mining across the manufacturer's internal quality management, supply chain, and CRM systems. This mapped the exact chronological footprint of the onboarding process across 40 early-adopter hospitals. The algorithm identified a severe, hidden bottleneck: legal redlining on data-sharing agreements was holding up operational training by an average of 35 days, a dependency that was entirely unnecessary.
Leadership decoupled the legal contracting workflow from the clinical training schedule and standardized the data-privacy language based on our analysis. The intervention cut average hospital onboarding time from 90 days to 45 days, rapidly expanding treatment capacity and allowing the therapy to hit its peak-year sales targets two full years ahead of the original forecast.