Substantiation for the choice of sample size calculation methods for effective monitoring of relevant pathogens within genomic epidemiological surveillance
- Authors: Gladkikh A.S.1, Naydenov D.D.1, Kotsar O.V.1, Belykh T.I.2, Dedkov V.G.1
-
Affiliations:
- Saint Petersburg Pasteur Institute
- Baikal State University
- Issue: Vol 103, No 3 (2026)
- Pages: 386-397
- Section: SCIENCE AND PRACTICE
- URL: https://microbiol.crie.ru/jour/article/view/19044
- DOI: https://doi.org/10.36233/0372-9311-801
- EDN: https://elibrary.ru/ZKDVGQ
- ID: 19044
Cite item
Abstract
Introduction. Continuous monitoring of the epidemiological situation is one of the key elements of a country's biological safety. The development and increasing accessibility of high-throughput sequencing technologies have led to the incorporation of genomic data into public health decision-making systems, known as genomic epidemiological surveillance. Effective implementation of genomic epidemiological surveillance for circulating pathogens requires finding a balance between data generation speed, estimation accuracy, and rational use of resources. Consequently, there is a need to develop evidence-based strategies for biological sample selection.
Aim: To justify the choice of methods for calculating the minimum sufficient sample size for sequencing depending on the objectives of epidemiological surveillance.
Materials and methods. This study presents a comparative analysis of two approaches to genomic monitoring: routine surveillance and targeted sentinel surveillance. The choice of statistical approaches for calculating the minimum sufficient sample size for each objective is justified, with specific examples that allow determining the required sequencing volume taking into account the current epidemiological situation and the prevalence of pathogen variants in the population.
Results. The study presents sample size calculation results depending on the objective: establishing the proportion of existing variants for routine surveillance, or detecting rare variants within sentinel surveillance at a predetermined prevalence threshold for a rare/new variant with a specified probability. For routine surveillance, the sample size was calculated based on the requirement for estimation precision of proportions (margin of error 5%) at a confidence level of 95%. For sentinel surveillance, the sample size was calculated under the condition that the probability of missing a variant (not detecting a single case) does not exceed 5%. The statistical results are presented in the form of graphs and calculation tables.
Conclusion. The proposed methods for calculating sample sizes for genomic monitoring depending on the objective can serve as a practical tool for epidemiological surveillance, increasing its efficiency. The use of these calculation methods will make it possible to obtain mathematically sound results regarding the genotypes of circulating pathogens and their degree of representation in the population.
Full Text
Introduction
Continuous monitoring of the epidemiological situation is one of the key elements of the biosecurity of a country. The COVID-19 pandemic, which clearly demonstrated the systemic limitations of national healthcare systems, has prompted a reevaluation of existing approaches to countering biological threats. The threat of new epidemics and the priority of their prevention necessitate a shift in epidemiological surveillance tactics through the use of modern analytical solutions.
The development and increased availability of high-throughput sequencing technologies have led to the integration of genomic data into the management decision-making process [1]. Genomic epidemiological surveillance, based on the systematic analysis of pathogen genetic diversity, is becoming a key tool for the early detection of new pathogen variants, the assessment of their evolutionary dynamics, and the monitoring of transmission routes and factors influencing virulence and antimicrobial resistance [2, 3]. However, even as the cost of sequencing has declined, this approach remains resource-intensive and involves significant time commitments, as well as the need to engage qualified personnel to collect and process data1. The scale of sequencing implementation worldwide remains highly uneven. In 2022, as part of the fight against the COVID-19 pandemic, genomic data for SARS-CoV-2 was available for 172 countries, but sequencing volumes ranged from a few sequences (e.g., 2 genomes in Vanuatu) to large-scale programs covering millions of samples (1.6 million in the United States) [4]. This heterogeneity reflects the fact that sampling volumes are often determined by the availability of laboratory capacity rather than actual epidemiological needs.
In response to the challenges outlined above, the World Health Organization (WHO) has developed the “Global Strategy for Genomic Surveillance of Pathogens with Pandemic and Epidemic Potential, 2022–2032.” [5]. The document outlines the main objectives and principles of the genomic epidemiological surveillance framework and identifies its five key tasks. The document aims to establish sequencing infrastructure in all WHO member states by 2032. However, the document is primarily conceptual in nature and requires further development at the national level, particularly the development of evidence-based algorithms to determine optimal study sizes for specific nosologies.
The effectiveness of genomic epidemiological surveillance is largely determined not only by the volume of data but also by the quality of the biological sample selection strategy. Determining the sample size is one of the most critically complex tasks and requires a balanced and scientifically sound approach [6]. The relevance of applying statistical approaches to sample size calculation increases in the context of the simultaneous circulation of multiple pathogen variants, the emergence of new genetic lineages, and the growing significance of interregional and cross-border spread of infections [3].
This article is devoted to justifying the choice of methods for calculating sample sizes depending on the objectives of genomic epidemiological surveillance. It examines the key principles for calculating sample size for whole-genome sequencing in the context of a specific epidemiological situation.
The aim of this study is to justify the selection of methods for calculating the minimum sufficient sample size for sequencing, depending on the objectives of epidemiological surveillance. When conducting genomic epidemiological surveillance, two fundamentally different tasks arise, each requiring distinct calculation methods:
- Routine surveillance: assuming that the data follow a binomial distribution (i.e., each sequenced sample either contains or does not contain a specific pathogen variant), determine the minimum sufficient sample size for accurately estimating the proportion of already known variants.
- Targeted surveillance: determine the sample size to ensure that the prevalence of a rare/new variant does not exceed a specified threshold.
The proposed approach allows the study size to be adapted to the current incidence rate (the size of the general population).
Materials and methods
Theoretical formulation of the objective
During routine surveillance, two distinct sub-tasks may be addressed, each of which requires the application of specific statistical approaches:
Routine surveillance sub-task 1: assessing the prevalence of a known pathogen variant. The objective is to determine the minimum sample size (n) that allows, with a given confidence level (1 – α, where α is the significance level, the acceptable probability of error in the estimate) and a margin of error (E), the estimation of the true proportion of the variant of interest in the population;
Routine surveillance sub-task 2: estimating the prevalence of 3 or more known variants of the pathogen. The objective is to determine the minimal sample size (n), that ensures the simultaneous estimation of the proportions of 3 or more circulating variants with a given confidence level (1 – α) and a margin of error (Е).
For the routine epidemiological surveillance sub-tasks under consideration, the following key parameters have been defined:
- population size (N) — the number of all newly diagnosed cases during a fixed observation period (e.g., one week);
- confidence probability (1 – α) — the probability that the constructed confidence interval covers the true value of the parameter being estimated (in this study, it is used as a calculation parameter when determining the required sample size). The calculations use the standard level 1 – α = 0.95 in biomedical research, which corresponds to the quantile of the standard normal distribution Z1 – α = 1,96;
- acceptable absolute error of the estimate (E) — half the width of the confidence interval for the estimated proportion of the target variant. The value E = 5% is commonly accepted in applied epidemiological studies, as it provides a compromise between the accuracy of statistical inference and economic feasibility, ensuring that the results are sufficiently informative for epidemiological surveillance tasks.
For routine surveillance purposes, the minimum sufficient sample size is calculated based on the approximation of the binomial distribution by the normal distribution, in accordance with the central limit theorem, provided that the conditions for the applicability of the approximation are met (np ≥ 10, n(1 – p) ≥ 10), where n is the calculated sample size, p is the proportion of the variant in the population.
In addition to routine assessment of the prevalence of known variants, a genomic epidemiological surveillance system may face the task of targeted surveillance—the early detection of new or rare genetic variants whose proportion in the population is still small. Unlike routine surveillance, the goal here is not to measure the exact proportion, but to verify the absence of a variant with a proportion above a specified threshold with a certain level of confidence. The goal is to calculate the minimum sample size (n) that allows for the detection of a new/rare variant with a given probability (1 − β), provided that its prevalence in the population is not below the threshold value.
Key parameters for targeted surveillance:
- the size of the population (N) — the number of newly confirmed cases during the reporting period (e.g., one week);
- the threshold proportion for a rare/new variant (q) — the proportion of cases in the population at which the variant should be detected by the surveillance system. The value of q is set by the researcher based on the epidemiological significance of the variant and the laboratory’s practical capabilities (e.g., 0.1%; 0.5%; or 1.0%);
- the expected number of carriers of the rare variant in the population (M) is calculated as the expected value of the binomial distribution M = Nq;
- β is the acceptable probability of failing to detect a rare variant; accordingly, (1 – β) specifies the probability of its successful detection. In this study, the standard level of β = 0.05 used in biomedical research is adopted, which corresponds to a 95% detection probability.
This study assumes that the sample is selected using simple random sampling from the population, and that the distribution of the characteristic (detection of the variant) follows a binomial distribution. The application of statistical formulas assumes the existence of an ideal test (100% sensitivity and specificity). This assumption allows us to focus on the main objective of the study — determining the minimum required sample size (n) that ensures a given level of detection probability. The final calculation is performed on the general population corresponding to the incidence rate in the region/state/country during the period under analysis. A correction factor for the laboratory is introduced to account for technical losses. Issues related to alternative sampling schemes, such as stratification by region, demographic characteristics, and other factors, are not addressed in this study. These aspects require separate analysis and fall outside the scope of the study’s objectives. The proposed methodology is a pilot study and is intended for further testing in healthcare practice.
Routine surveillance, subtask 1: Estimating the proportion of a known variant in the population
For an infinite or very large population, the basic (minimum sufficient) sample size is calculated based on the required confidence level and precision of the estimate, using the normal approximation of the binomial distribution of the sampling proportion [7]:
, (1)
where Z is the quantile of the normal distribution for a chosen confidence level of 95%; p is the estimated proportion of the target option; and E is the margin of error, set to 0.05.
In the absence of prior information about the proportion of the option, the value p = 0.5 is assumed, which maximizes the variance of the sampling distribution and, as a result, provides the most conservative estimate of the required sample size, ensuring that the specified requirements for precision and confidence level are met for any true value of the proportion. To account for the limited population size, the standard correction for the finite population is used [8]:
, (2)
where N is the number of newly diagnosed cases during the study period; n is the adjusted sample size.
Calculation example. Suppose N = 1,000 cases of the disease were registered over the course of a week. Assume that variant X is present in approximately 60% (p = 0.6; if p is unknown, p = 0.5 is always used) of those who fell ill. We need to estimate, with 95% confidence (Z = 1.96) and a margin of error of 5% (E = 0.05), what proportion of the population is accounted for by variant X.
Using formula (1) with the selected confidence level and margin of error, the base sample size will be equal to:
Since the calculated sample size represents a significant proportion of the population size (N), we apply a correction for the target population using formula (2):
Thus, to estimate the proportion of variant X with p = 0.6 among 1,000 cases with an accuracy of ±5% and a 95% confidence level, 270 samples must be analyzed, which represents 27% of all samples (270/1,000).
Routine surveillance, subtask 2: Estimating the proportion of each pathogen variant in the population (when there are three or more variants)
When three or more variants of a pathogen (m ≥ 3) are circulating simultaneously, the task of quantifying the population structure boils down to estimating multinomial proportions (category shares). It is necessary to determine the sample size n that ensures a given margin of error E for all estimated proportions simultaneously at a confidence level of (1 – α).
This study employs a conservative approach proposed by R.D. Tortora [9], which guarantees the required precision under any distribution of proportions (the so-called “worst-case” scenario). Tortora’s formula is based on the χ2 quantile of a distribution with one degree of freedom and uses the Bonferroni correction for multiple comparisons, which allows for the construction of simultaneous confidence intervals for multinomial proportions. This method is computationally simple, making it convenient for practical use:
, (3)
where m is the number of categories (options) being estimated simultaneously, and is the quantile of the distribution with one degree of freedom at the chosen confidence level 1 – α = 0,95 and margin of error E = 0,05.
The resulting value of the base sample size n0 is then adjusted to account for the final population size N (Formula 2).
Calculation example. Example of calculating the adjusted confidence probability for each category (with the Bonferroni correction) given an overall confidence probability 1 – α = 0.95 and number of options m = 3:
Given the adjusted confidence level, the desired quantile of the χ2 distribution is χ20.98333;1 ≈ 5.73. This value can be obtained using statistical tables, programming languages, or software (R, Python, MS Excel). Appendix 1 on the journal’s website provides a table of quantiles for the chi-square distribution with one degree of freedom for multiple comparisons when the number of circulating variants ranges from 3 to 10.
Calculation example. Suppose that N = 500 cases of the disease were recorded over the course of a week. It is necessary to estimate the proportions of m = 3 pathogen variants with a margin of error E = 0.05 and a confidence probability 1 – α = 0.95. To ensure the simultaneous accuracy of all estimates, we apply the Bonferroni correction. The corrected quantile of the χ² distribution with one degree of freedom is ≈ 5.73. Then the base sample size n0 is calculated using formula (3)
Next, we apply formula (2) to calculate the sample size, taking into account the incidence rate N, and obtain:
Thus, to simultaneously estimate the proportions of the three circulating variants among 500 newly reported cases, with a margin of error of 5% and a 95% confidence level, the required sample size is 268 samples, which represents 53.6% of all samples (268/500).
Targeted screening: identifying new or rare genetic variants that are currently present in only a small proportion of the population
Unlike routine surveillance, where the primary goal is to estimate the frequencies of known variants with a specified detection limit, the objective of targeted surveillance is the early detection of new or rare genetic variants whose prevalence in the population is still low. Here, it is necessary to determine the minimum sample size n that ensures a given probability of detecting at least one instance of a rare variant, provided that its prevalence in the general population is not lower than a predefined threshold.
The threshold is set by the researcher based on the variant’s epidemiological significance, as well as the laboratory’s technical and financial capacity to perform the required amount of sequencing. The lower the prevalence threshold, the larger the sample size required for sequencing; therefore, the choice of threshold represents a trade-off between the desired sensitivity of the epidemiological surveillance system and available resources. In practice, thresholds ranging from 0.1% to 1% are typically used, depending on the surveillance objectives and the epidemiological situation.
The probability of failing to detect a rare variant (P0) in a sample of size n drawn from a population of size N containing M carriers of that variant is described by the hypergeometric distribution [10].
M is calculated based on the prevalence threshold selected by the researcher: M = Nq , where q is the threshold proportion of the rare variant in the population.
, (4)
where N is the number of newly diagnosed cases during the study period (the general population); M = Nq is the estimated number of carriers of the rare variant in the general population; q is the threshold prevalence of the rare variant in the general population (e.g., 0.001; 0.005; 0.01) depending on the preset threshold value (e.g., 0.1%; 0.5%; 1%); n is the sample size (number of tested samples); i is the iteration number.
Given the computational complexity of Equation (4), it is advisable to use the Cannon and Roy approximation [11]. When the variant has low prevalence in the population, the assumption P0 ≤ β allows us, through appropriate algebraic transformations, to derive a formula for approximating the minimum required sample size n:
, (5)
where N is the number of newly diagnosed cases during the study period; M = Nq is the estimated number of carriers of the rare variant in the general population; q is the threshold prevalence of the rare variant in the general population (e.g., 0.001; 0.005; 0.01) depending on the preset threshold value (e.g., 0.1%; 0.5%; 1%); n — sample size; β — acceptable probability of failing to detect the rare variant (e.g., β = 0.05); (1 – β) — probability of detecting at least 1 case of the rare variant, power of the method (e.g., 0.95).
Calculation example. Suppose that N = 10,000 cases of the disease were registered over the course of a week. It is assumed that a new variant of pathogen Y may be present in the population with a prevalence threshold of 0.5%. The expected number of carriers of variant Y in the registered population of cases N is calculated. The required sample size n is determined so that, with a 95% probability of detection, at least 1 case of the new variant Y is identified if its prevalence in the population is not lower than the specified threshold of 0.5%.
Thus, N = 10,000, q = 0.005, β = 0.05. Then the expected number of carriers of pathogen variant Y in the registered population of cases is:
Substituting the resulting value into equation (5), we get:
Thus, among 10,000 cases, approximately 50 carriers of pathogen variant Y are expected. To ensure 95% statistical power, a minimum of 581 samples must be tested.
If the new Y variant is not detected after testing the calculated number of samples, this provides grounds to conclude that, under the given assumptions, the prevalence of the variant does not exceed the specified threshold. If the Y variant is detected, further quantitative assessment of its prevalence in the population should be performed as part of subtask 1 of routine monitoring.
Laboratory efficiency correction factor
In all of the scenarios considered, it is further proposed to account for the proportion of samples that successfully pass quality control (QC) during the sample preparation and sequencing stages. To this end, a laboratory efficiency correction factor kqc is introduced, which takes values in the range 0 < kqc < 1 [12].
The actual order volume for nfct sequencing is calculated as follows:ом:
, (6)
where n is the theoretically calculated statistically required sample size, and kqc is the empirically determined proportion of successfully sequenced samples in a specific laboratory (e.g., kqc = 0.9).
Thus, with a technical loss rate of 10% (kqc = 0.9), the actual number of samples sent for sequencing must be increased by a factor of 1/0.9 ≈ 1.11 compared to the theoretically calculated volume.
Statistical analysis and visualization
Statistical calculations aimed at estimating the required sequencing sample size in various scenarios of pathogen circulation were performed using the Python programming language, version 3.12.12. Mathematical calculations were performed using the NumPy library. Quantiles of the χ² distribution for multinomial proportions were calculated using the scipy.stats module of the SciPy library (v. 1.17.0). Tabular data processing and aggregation were performed in the Pandas environment. Visualization of the simulation results was implemented using the Matplotlib and Seaborn libraries. When plotting graphs of the relationship between sample size and population size, a logarithmic scale was used on the x-axis, which allowed for the correct representation of the effect of statistical saturation across a wide range of population sizes N.
Results
Sample size dynamics in the supervision of a known variant (routine supervision, subtask 1)
The results regarding sample size demonstrated the presence of a pronounced statistical saturation effect (plateau). As the population size N increases, the required sample size n grows nonlinearly and converges quite rapidly to an asymptotic value, i.e., it stabilizes (Table 1, Fig. 1). When the incidence exceeds 50,000 cases (N > 50,000), the relationship reaches a plateau with an asymptotic value of approximately 384 samples. A further increase in the number of cases does not require a proportional increase in sequencing volume while maintaining the same statistical accuracy.
Table 1. Calculation of the sample size for estimating the allele frequency in populations under conservative assumptions (p = 0.5; 95% confidence level; 5% margin of error)
Number of confirmed cases (N) | Estimated sample size for the study (n) | % of confirmed cases | Number of samples for the study, adjusted for laboratory performance (kqc = 0.9) |
100 | 80 | 80.0 | 89 |
500 | 218 | 43.6 | 242 |
1000 | 278 | 27.8 | 309 |
5000 | 357 | 7.1 | 397 |
10,000 | 370 | 3.7 | 411 |
20,000 | 377 | 1.9 | 419 |
50,000 | 381 | 0.8 | 423 |
> 100,000 (plateau) | 384 | < 0.4 | 427 |
Fig. 1. The relationship between the required minimum sample size (n) and the number of newly confirmed cases (N) for estimating the proportion of a specific pathogen variant in the population, with a 95% confidence level and a 5% margin of error.
The x-axis is on a logarithmic scale; the effect of statistical saturation (plateau) is shown.
Trends in sample size for monitoring the circulation of three or more known variants (routine surveillance, subtask 2)
When three or more variants of a pathogen are circulating simultaneously, the required sample sizes increase significantly compared to the task of estimating the proportion of a single variant. The calculation results presented in Table 2 and Fig. 2 demonstrate that, with high prevalence, the calculated sample size stabilizes at approximately 570 samples, which is about 1.5 times higher than the value obtained for subtask 1. This is due to the need to ensure the specified precision (E = 0.05) for all estimated proportions simultaneously while maintaining the overall confidence level (1 – α) = 0.95. As in the previous case, the relationship between sample size and population size is nonlinear with a pronounced effect of statistical saturation: when N > 50,000, a further increase in incidence requires virtually no increase in sample size.
Table 2. Comparison of estimated sample sizes for estimating the proportions of variants 2 and 3 circulating simultaneously in the population (implementation of subtasks 1 and 2 of routine surveillance)
Number of confirmed cases | 2 variants in the population | 3 variants in the population | Growth rate |
100 | 80 | 85 | 1.06 |
500 | 218 | 268 | 1.23 |
1000 | 278 | 365 | 1.31 |
5000 | 357 | 515 | 1.44 |
10,000 | 370 | 543 | 1.47 |
20,000 | 377 | 557 | 1.48 |
50,000 | 381 | 567 | 1.48 |
> 100,000 (plateau) | 384 | 570 | 1.49 |
Fig. 2. The relationship between the required minimum sample size (n) and the number of newly confirmed cases (N) when two and three variants of the pathogen are circulating simultaneously.
The x-axis is on a logarithmic scale; the effect of statistical saturation (plateau) is shown.
Changes in sample size during targeted surveillance
Unlike routine surveillance, the objectives of targeted surveillance are focused on the early detection of new or rare variants that account for a small proportion of the population. The results of the simulation of surveillance monitoring are presented in Table 3 and Fig. 3. They demonstrate a marked dependence of the required sample size on the specified prevalence threshold: the lower the threshold, the more samples are required to ensure a specified detection probability of 1 − β = 0.95.
Table 3. Sample size required to detect a rare variant with 95% confidence at various prevalence thresholds (for N > 10,000)
The prevalence threshold for a rare variant, % | Prevalence | Required sample size (n) | Number of samples for the study, adjusted for laboratory performance (kqc = 0.9) |
5.0 | 1 to 20 | 59 | 66 |
2.0 | 1 to 50 | 149 | 166 |
1.0 | 1 to 100 | 298 | 331 |
0.5 | 1 to 200 | 596 | 662 |
0.1 | 1 to 1000 | 2950 | 3278 |
Note. The calculation in the table is based on N = 100,000; taking into account the saturation effect of the curve, it can serve as a rough guide for N > 10,000.
Fig. 3. The relationship between the required minimum sample size (n) and the number of newly confirmed cases (N) for various threshold values of the prevalence of a rare/new variant (95% detection probability).
The x-axis is on a logarithmic scale.
The values presented in Table 3 were obtained using an approximate formula for the hypergeometric model (Formula 5), assuming that the population size is sufficiently large (N > 10,000). To detect a variant with a prevalence threshold of 0.1%, approximately 3,000 samples are required, which is more than seven times the sample size needed for routine assessment of the frequency of a known variant. The graph in Fig. 3 shows that, for a fixed prevalence threshold, the required sample size quickly reaches a plateau as the number of newly confirmed cases (N) increases. Therefore, the values given in Table 3 should be considered approximate for high prevalence rates, whereas exact calculations for specific values of N are presented in Appendix 2.
If the target variant is not detected in a sample of size n, this provides grounds to conclude that its prevalence in the population does not exceed a specified prevalence threshold; in this case, the specified detection probability is 1 − β = 0.95. If the variant is detected, further quantitative estimation of its proportion should be performed using routine surveillance methods (subtask 1) aimed at estimating parameters with a specified margin of error.
Thus, the sentinel surveillance approach differs fundamentally from routine surveillance, as it is not aimed at comprehensively recording all cases of a variant’s circulation, but rather at the early detection of its emergence based on predefined detection probabilities and prevalence thresholds. Sentinel surveillance should be viewed as an adaptive tool that allows for dynamic changes in the sample size for sequencing depending on the epidemiological situation and interim analysis results.
Discussion
Although genomic sequencing has become widely adopted in many biological and medical laboratories over the past few decades, it still requires significant resources, including financial, infrastructural, and human resources. Given limited resources and the need for rapid decision-making during outbreaks and epidemics, there is a need to reduce sample sizes as much as possible while maintaining a reliable picture of the current epidemiological situation.
Genomic surveillance of SARS-CoV-2 during the COVID-19 pandemic has become a key tool in the epidemiologists’ arsenal. Its use has made it possible to track the emergence and spread of new variants of the virus, analyze its evolutionary dynamics, and adjust response strategies in a timely manner. The adopted “Global genomic surveillance strategy for pathogens with pandemic and epidemic potential, 2022–2032” [5] postulated the necessity for genomic monitoring of all infectious diseases with epidemic potential. To date, genomic epidemiological surveillance has become an indispensable tool in the fight against epidemics in Russia [13–15]; however, there are still no uniform algorithms or approaches for its implementation. In many countries, including Russia, large-scale sequencing initiatives are often designed primarily based on technical capabilities rather than scientific or statistical rationale [3, 16].
The approaches proposed in this study clearly demonstrate that when designing samples for genomic sequencing, it is necessary to consider that there are at least two fundamentally different epidemiological tasks requiring different analytical approaches: routine surveillance and targeted surveillance. Routine surveillance involves a regular but relatively small volume of samples. Targeted surveillance is focused on identifying rare variants and requires a significantly larger number of tests, especially when a low prevalence threshold (0.1%) is set. Recommended approximate sample sizes for sequencing at given incidence ranges, depending on the objective of genomic epidemiological surveillance, are provided in Appendix 3.
Consequently, the choice of approach should be determined not merely by the technical capabilities of the laboratories, but by research and practical objectives. Depending on the task at hand, the genomic monitoring strategy must be tailored. Using a whole-genome sequencing strategy or a fixed percentage of all positive samples proves to be economically inefficient and statistically unjustified during periods of high incidence.
The identified statistical plateau effect is of fundamental importance for managing healthcare system resources. During periods of worsening epidemiological conditions, laboratories may not need to increase sequencing volumes in proportion to the rise in incidence; at the same time, the calculated sample size remains sufficient to ensure the informativeness of results required for genomic epidemiological surveillance. It is advisable to allocate the freed-up resources to targeted sequencing of high-risk groups, severe cases, and atypical cases.
Conclusion
- There is no universal approach to designing sequencing samples. A sample of no more than 384 specimens is sufficient to estimate the prevalence of a known pathogen variant (for N → ∞), whereas samples of more than 2,500 specimens are required to detect rare or novel variants at a 0.1% threshold.
- When incidence rates are high during a specific time period (with more than 10,000 newly confirmed cases), further increases in incidence have virtually no effect on the required sample size.
- It is recommended to implement dynamic quota allocation, where the sample size is adjusted weekly based on the current epidemiological situation, the task at hand (detection of a new pathogen variant or monitoring of already known variants), and the laboratory coefficient (kqc).
Thus, the proposed methods for calculating sample sizes for genomic surveillance, depending on the specific task at hand, can serve as a reliable tool for epidemiological surveillance, thereby enhancing its effectiveness. The use of calculation formulas and tables will prevent laboratories from becoming overburdened, while ensuring mathematically sound calculations of sequencing volumes sufficient to assess the genetic diversity of circulating pathogens and the extent of their prevalence in the population.
1 World Health Organization. SARS-CoV-2 genomic sequencing for public health goals: interim guidance. 2021. URL: https://www.who.int/publications/i/item/WHO-2019-nCoV-genomic_sequencing-2021.1 (data of access: 27.01.2026).
About the authors
Anna S. Gladkikh
Saint Petersburg Pasteur Institute
Author for correspondence.
Email: angladkikh@gmail.com
ORCID iD: 0000-0001-6759-1907
Cand. Sci. (Biol.), senior researcher, Laboratory of molecular genetic monitoring
Russian Federation, St. PetersburgDmitry D. Naydenov
Saint Petersburg Pasteur Institute
Email: dmitnayd@gmail.com
ORCID iD: 0009-0009-2683-084X
junior researcher, Laboratory of molecular genetic monitoring
Russian Federation, St. PetersburgOleg V. Kotsar
Saint Petersburg Pasteur Institute
Email: kotsar@pasteurorg.ru
ORCID iD: 0009-0005-7291-5167
researcher, Laboratory of biomedical statistics
Russian Federation, St. PetersburgTatiana I. Belykh
Baikal State University
Email: belyhti@bgu.ru
ORCID iD: 0009-0009-2488-3630
Cand. Sci. (Physics and Mathematics), Associate Professor, Department of mathematical methods and digital technologies
Russian Federation, IrkutskVladimir G. Dedkov
Saint Petersburg Pasteur Institute
Email: vgdedkov@yandex.ru
ORCID iD: 0000-0002-5500-0169
Cand. Sci. (Med.), Deputy director for science
Russian Federation, St. PetersburgReferences
- Barbian H.J., Kittner A., Teran R., et al. A response playbook for early detection and population surveillance of new SARS-CoV-2 variants in a regional public health laboratory. BMC Public Health. 2024;24(1):59. DOI: https://doi.org/10.1186/s12889-023-17536-0
- Hicks A.L., Kissler S.M., Mortimer T.D., et al. Targeted surveillance strategies for efficient detection of novel antibiotic resistance variants. Elife. 2020;9:e56367. DOI: https://doi.org/10.7554/elife.56367
- Ling-Hu T., Rios-Guzman E., Lorenzo-Redondo R., et al. Challenges and opportunities for global genomic surveillance strategies in the COVID-19 era. Viruses. 2022;14(11):2532. DOI: https://doi.org/10.3390/v14112532
- Chen Z., Azman A.S., Chen X., et al. Global landscape of SARS-CoV-2 genomic surveillance and data sharing. Nat. Genet. 2022;54(4):499–507. DOI: https://doi.org/10.1038/s41588-022-01033-y
- ВОЗ. Глобальная стратегия геномного эпиднадзора за возбудителями болезней, обладающих пандемическим и эпидемическим потенциалом, 2022–2032 гг;2022. Available at: https://who.int/ru/publications/i/item/9789240046979 WHO. Global genomic surveillance strategy for pathogens with pandemic and epidemic potential, 2022–2032;2022. Available at: https://who.int/publications/i/item/9789240046979
- Contreras S., Oróstica K.Y., Daza-Sanchez A., et al. Model-based assessment of sampling protocols for infectious disease genomic surveillance. Chaos Soliton Fract. 2023;167:113093. DOI: https://doi.org/10.1016/j.chaos.2022.113093
- Naing L., Winn T., Rusli B.N. Practical issues in calculating the sample size for prevalence studies. Arch. Orofac. Sci. 2006;1:9–14.
- Rodríguez Del Águila M., González-Ramírez A. Sample size calculation. Allergol. Immunopathol. (Madr.). 2014;42(5):485–92. DOI: https://doi.org/10.1016/j.aller.2013.03.008
- Tortora R.D. A note on sample size estimation for multinomial populations. Am. Stat. 1978;32(3):100–2. DOI: https://doi.org/10.2307/2683352
- Cameron A.R., Baldock F.C. A new probability formula for surveys to substantiate freedom from disease. Prev. Vet. Med. 1998;34(1):1–17. DOI: https://doi.org/10.1016/s0167-5877(97)00081-0
- Cannon R.M., Roe R.T. Livestock Disease Surveys: A Field Manual for Veterinarians. Canberra;1982.
- Bolarinwa O.A. Sample size estimation for health and social science researchers: The principles and considerations for different study designs. Niger. Postgrad. Med. J. 2020;27(2):67–75. DOI: https://doi.org/10.4103/npmj.npmj_19_20
- Акимкин В.Г., Семененко Т.А., Хафизов К.Ф. и др. Биобезопасность и геномный эпидемиологический надзор. Эпидемиология и вакцинопрофилактика. 2024;23(5):4–12. Akimkin V.G., Semenenko T.A., Khafizov K.F., et al. Biosafety and genomic epidemiological surveillance. Epidemiology and Vaccinal Prevention. 2024;23(5):4–12. DOI: https://doi.org/10.31631/2073-3046-2024-23-5-4-12 EDN: https://elibrary.ru/jjrtnk
- Акимкин В.Г., Семененко Т.А., Хафизов К.Ф. и др. Стратегия геномного эпидемиологического надзора. Проблемы и перспективы. Журнал микробиологии, эпидемиологии и иммунобиологии. 2024;101(2):163–72. Akimkin V.G., Semenenko T.A., Khafizov K.F., et al. Genomic surveillance strategy. Problems and perspectives. Journal of Microbiology, Epidemiology and Immunobiology. 2024;101(2):163–72. DOI: https://doi.org/10.36233/0372-9311-507 EDN: https://elibrary.ru/mymnik
- Котов И.А., Аглетдинов М.Р., Роев Г.В. и др. Геномный надзор за SARS-CoV-2 в Российской Федерации: возможности платформы VGARus. Журнал микробиологии, эпидемиологии и иммунобиологии. 2024;101(4):435–47. Kotov I.A., Agletdinov M.R., Roev G.V., et al. Genomic surveillance of SARS-CoV-2 in Russia: insights from the VGARus platform. Journal of Microbiology, Epidemiology and Immunobiology. 2024;101(4):435–47. DOI: https://doi.org/10.36233/0372-9311-554 EDN: https://elibrary.ru/irjxcx
- Gangavarapu K., Latif A.A., Mullen J.L., et al. Outbreak.info genomic reports: scalable and dynamic surveillance of SARS-CoV-2 variants and mutations. Nat. Methods. 2023;20(4):512–22. DOI: https://doi.org/10.1038/s41592-023-01769-3
Supplementary files






