MODELLING HEALTH INSURANCE CLAIMS USING STATISTICAL METHODS
Get complete chapters, abstract, references and questionnaire delivered to your WhatsApp or email.
CHAPTER ONE
INTRODUCTION
1.1 Background to the Study
Health insurance operates on the principle of risk pooling,
in which the uncertain and potentially catastrophic medical expenditure of an
individual is transferred to an insurer in exchange for a comparatively small
and certain premium. The financial viability of any such arrangement depends
almost entirely on the insurer's ability to predict, with acceptable accuracy,
how often claims will arise and how large they will be. This predictive
exercise is fundamentally statistical: claim occurrence is a counting process,
claim size is a continuous and typically right-skewed random variable, and the
aggregate liability of the insurer is a compound distribution formed from the
two (Klugman, Panjer, & Willmot, 2019).
The classical actuarial approach decomposes aggregate claims
into a frequency component and a severity component. Frequency is
conventionally modelled using the Poisson, negative binomial, or zero-inflated
count distributions, while severity is modelled using the gamma, lognormal,
Weibull, or Pareto families. Generalised Linear Models (GLMs) extended this
framework by allowing rating factors such as age, sex, plan type, and provider
category to be incorporated as covariates. A long-standing simplifying
assumption in this framework is that frequency and severity are independent, an
assumption that is theoretically convenient because it permits the two
components to be estimated separately. Empirical evidence has, however,
repeatedly contradicted it. Frees, Gao and Rosenberg (2011) demonstrated
dependence between the number and size of health-care expenditure claims, while
Erhardt and Czado (2012) modelled dependent yearly claim totals, including zero
claims, in private health insurance. Shi, Feng and Ivantsova (2015) formalised
a two-part dependent frequency severity framework, and more recent
contributions have relaxed distributional and functional-form assumptions
altogether through machine-learning estimators such as gradient boosting and
copula-based hierarchical models (Power, Côté, & Duchesne, 2024; So &
Valdez, 2024; Xu, Manathunga, & Hong, 2023).
In Nigeria, the statistical modelling of health insurance
claims has assumed policy urgency. The National Health Insurance Authority
(NHIA) Act, 2022, which repealed the National Health Insurance Scheme Act of
2004, made health insurance mandatory for all residents and created a
Vulnerable Group Fund, thereby signalling a decisive shift towards universal
health coverage (Adewole & Adebowale, 2022; Ipinnimo, Durowade, Afolayan,
Ajayi, & Akande, 2022). The reform was driven by a persistent coverage
failure: after more than two decades of operation, the old scheme had enrolled
less than ten per cent of the population, and survey evidence indicates that
only about two to four per cent of Nigerian adults reported any health
insurance cover in recent national surveys (Aloh, Onwujekwe, Aloh, &
Nwankwo, 2020; Okoroafor et al., 2025). As enrolment expands under the new
mandate, Health Maintenance Organisations (HMOs) and the Authority itself will
confront claim volumes and claim-cost distributions for which there is little
accumulated Nigerian actuarial experience.
This matters because the pricing of capitation and
fee-for-service arrangements, the setting of provider tariffs, the estimation
of technical provisions, and the detection of fraudulent or inflated claims all
rest on distributional assumptions about the claims process. Where those
assumptions are imported wholesale from foreign experience or set
administratively rather than estimated from data, the result is either
under-pricing, which threatens solvency, or over-pricing, which suppresses
enrolment. Fraudulent and exaggerated claims have already been identified as a
material drain on Nigeria's health financing architecture (Ogunbekun &
Adeleye, 2023). A rigorous, data-driven model of the Nigerian health insurance
claims process is therefore not an academic luxury but a precondition for the
sustainability of the reform.
This study responds to that need by fitting, comparing, and
validating a set of candidate statistical models for claim frequency and claim
severity on Nigerian health insurance claims data, and by using the selected
models to estimate aggregate claim distributions and pure premiums.
1.2 Statement of the Problem
Despite the centrality of claims modelling to health
insurance solvency, Nigerian insurers and HMOs largely set premiums and
capitation rates on the basis of administrative negotiation, historical
averages, and competitor benchmarking rather than fitted probability models.
Three problems follow.
First, the use of simple arithmetic averages ignores the
heavy right tail characteristic of medical claim severity, so that the
probability of extreme aggregate loss is systematically understated. Second,
the conventional assumption of independence between claim frequency and claim
severity, still embedded in most practical rating exercises, has been shown to
be empirically false in health portfolios (Erhardt & Czado, 2012; Frees et
al., 2011), leading to biased estimates of aggregate loss. Third, the Nigerian
literature on health insurance is dominated by studies of enrolment, awareness,
willingness to pay, and policy evaluation, with comparatively little
quantitative modelling of the claims process itself using local data.
The consequence is a knowledge gap: it is not currently
established which probability distributions best describe the frequency and
severity of health insurance claims in Nigeria, whether dependence exists
between the two components, or how much pricing error results from ignoring it.
Without this evidence, the expansion of coverage envisaged by the NHIA Act,
2022 risks being built on actuarially unsound foundations. This study addresses
that gap.
1.3 Aim and Objectives of the Study
The aim of this study is to model health insurance claims
using appropriate statistical methods in order to improve the estimation of
claim liabilities and premium rates.
The specific objectives are to:
1.
examine
the distributional characteristics of health insurance claim frequency and
claim severity data;
2.
fit
and compare competing count models (Poisson, negative binomial and
zero-inflated variants) to claim frequency data;
3.
fit
and compare competing severity models (gamma, lognormal, Weibull and Pareto) to
claim amount data;
4.
determine
whether a statistically significant dependence exists between claim frequency
and claim severity;
5.
estimate
the aggregate claim distribution and the pure premium using the best-fitting
frequency and severity models; and
6.
assess
the predictive accuracy of the selected models using appropriate validation
criteria.
1.4 Research Questions
1.
What
are the distributional characteristics of health insurance claim frequency and
severity in the study portfolio?
2.
Which
count distribution provides the best fit to observed claim frequency?
3.
Which
continuous distribution provides the best fit to observed claim severity?
4.
Is
there a statistically significant dependence between claim frequency and claim
severity?
5.
What
aggregate claim distribution and pure premium result from the fitted models?
6.
How
accurately do the selected models predict out-of-sample claim experience?
1.5 Research Hypotheses
The following null hypotheses will be tested at the 5% level
of significance:
H₀₁: The Poisson distribution does not
provide a significantly better fit to health insurance claim frequency than the
negative binomial distribution.
H₀₂: There is no significant difference in
goodness of fit among the gamma, lognormal, Weibull and Pareto distributions
for claim severity.
H₀₃: There is no significant dependence
between claim frequency and claim severity.
H₀₄: Claimant characteristics (age, sex,
plan type and provider category) have no significant effect on expected claim
cost.
1.6 Significance of the Study
The study will benefit several groups. For insurers
and HMOs, it provides an empirically grounded basis for premium
loading, capitation rate-setting and reserve estimation, replacing
rule-of-thumb pricing with fitted models. For the National Health
Insurance Authority and state health insurance agencies, it supplies
evidence on the cost structure of claims that can inform benefit-package design
and the sizing of the Vulnerable Group Fund. For regulators (NAICOM and
the NHIA), the fitted distributions offer a benchmark against which
submitted technical provisions can be reviewed, which is directly relevant
under the risk-based capital regime introduced by the Nigerian Insurance
Industry Reform Act, 2025. For scholarship, the study extends
the international frequency severity literature to a Nigerian data setting that
has been largely unexamined, and tests in that setting the dependence findings
reported by Frees et al. (2011) and Erhardt and Czado (2012). Finally, for policyholders,
more accurate pricing reduces the risk of insurer insolvency and of the
arbitrary premium increases that follow mis-priced portfolios.
1.7 Scope of the Study
The study covers health insurance claims data obtained from a
selected Health Maintenance Organisation (or set of HMOs) operating in Nigeria
over a period of not less than five consecutive years. The variables of
interest are the number of claims per enrollee per period, the amount paid per
claim, and the enrollee and plan characteristics available in the
administrative records. Methodologically, the study is limited to parametric
and semi-parametric statistical models of frequency, severity and aggregate
loss; it does not extend to network pricing, provider contracting design, or
health outcome evaluation.
1.8 Limitations of the Study
The study is constrained by the following. (i) Data
access and confidentiality claims data are commercially sensitive, and
insurers may release only anonymised or partially aggregated extracts, limiting
the covariates available for modelling. (ii) Data quality Nigerian administrative claims records are
known to contain coding inconsistencies, missing entries and duplicate
submissions, which require cleaning that may itself alter distributional
properties. (iii) Reporting delay claims
incurred but not reported (IBNR) at the data cut-off are unobserved, so
estimates of the most recent period may be understated. (iv) Generalisability
results obtained from one or a few HMOs may
not describe the whole Nigerian market. (v) Inflationary distortion
the high and volatile medical-cost inflation
experienced in Nigeria during the study period complicates comparison of
nominal claim amounts across years, and adjustment assumptions will be
required.
1.9 Operational Definition of Terms
Claim frequency: The number of claims
arising from a policy or enrollee within a defined period of exposure.
Claim severity: The monetary amount of an
individual claim.
Aggregate claim: The total of all claim
amounts arising from a portfolio within a period, modelled as a compound
distribution of frequency and severity.
Pure premium: The expected claim cost per
unit of exposure, exclusive of expenses, profit and contingency loadings.
Capitation: A fixed periodic payment made to
a health-care provider per enrolled person, irrespective of the volume of
services rendered.
Health Maintenance Organisation (HMO): An
organisation accredited by the NHIA to manage the provision of health-care
services to enrollees under a health insurance plan.
Zero-inflation: The presence in count data
of more zero observations than the assumed count distribution predicts.
Goodness of fit: The degree of agreement
between a fitted theoretical distribution and the observed empirical
distribution, assessed here by chi-square, Kolmogorov Smirnov, Anderson Darling
and information-criterion statistics.
References
Adewole, D. A., & Adebowale, A. S. (2022). The new
National Health Insurance Act of Nigeria: How it will insure the poor and
ensure universal health coverage. Population Medicine, 4(December),
33.
Aloh, H. E., Onwujekwe, O. E., Aloh, O. G., & Nwankwo, C.
H. (2020). Is bed turnover rate a good metric for hospital scale efficiency? A
measure of resource utilization rate for hospitals in Southeast Nigeria. Cost
Effectiveness and Resource Allocation, 18(1), 21. https://doi.org/10.1186/s12962-020-00216-w
Erhardt, V., & Czado, C. (2012). Modeling dependent
yearly claim totals including zero claims in private health insurance. Scandinavian
Actuarial Journal, 2012(2), 106 129. https://doi.org/10.1080/03461238.2010.489762
Frees, E. W., Gao, J., & Rosenberg, M. A. (2011).
Predicting the frequency and amount of health care expenditures. North
American Actuarial Journal, 15(3), 377 392. https://doi.org/10.1080/10920277.2011.10597626
Gschlößl, S., & Czado, C. (2007). Spatial modelling of
claim frequency and claim size in non-life insurance. Scandinavian
Actuarial Journal, 2007(3), 202 225. https://doi.org/10.1080/03461230701414764
Ipinnimo, T. M., Durowade, K. A., Afolayan, C. A., Ajayi, P.
O., & Akande, T. M. (2022). The Nigeria National Health Insurance Authority
Act and its implications towards achieving universal health coverage. Nigerian
Postgraduate Medical Journal, 29(4), 281 287. https://doi.org/10.4103/npmj.npmj_216_22
Klugman, S. A., Panjer, H. H., & Willmot, G. E. (2019). Loss
models: From data to decisions (5th ed.). John Wiley & Sons.
National Health Insurance Authority. (2022). National
Health Insurance Authority Act, 2022. Federal Government of Nigeria.
Okoroafor, S. C., Ahmat, A., Asamani, J. A., Millogo, J. J.
S., & Nyoni, J. (2025). Health insurance in Nigeria: Findings from the
People's Voice Survey. PLOS Global Public Health, 5(9), e0006948. https://doi.org/10.1371/journal.pgph.0006948
Power, J., Côté, M.-P., & Duchesne, T. (2024). A flexible
hierarchical insurance claims model with gradient boosting and copulas. North
American Actuarial Journal, 28(4), 772 800. https://doi.org/10.1080/10920277.2023.2279782
Shi, P., Feng, X., & Ivantsova, A. (2015). Dependent
frequency severity modeling of insurance claims. Insurance: Mathematics and
Economics, 64, 417 428. https://doi.org/10.1016/j.insmatheco.2015.07.006
So, B., & Valdez, E. A. (2024). Zero-inflated Tweedie
boosted trees with CatBoost for insurance loss analytics
(arXiv:2406.16206). arXiv. https://doi.org/10.48550/arXiv.2406.16206
Xu, S., Manathunga, V., & Hong, D. (2023). Framework of
BERT-based NLP models for frequency and severity in insurance claims. Variance,
16(2). https://doi.org/10.66573/001c.89002
This project contains full academic material including literature review, methodology,
data analysis and conclusion.
VERIFIED COMPLETE RESEARCH PROJECT TOPICS AND MATERIALS
58 PAGES
Need a Custom Project Written for You?
Our professional writers can write a unique, plagiarism-free project on any topic in your department — delivered before your deadline.