
When a dangerous disease starts spreading, public health teams need fast, reliable answers. How quickly is the outbreak growing? How many hospital beds will be needed next week? Is the outbreak slowing down or speeding up in different geographic areas of the north-eastern DRC? This page explains the analysis workflow that is being developed by the Jameel Institute & the Centre for Global Infectious Disease Analysis at Imperial, together with its partners from the DRC and elsewhere. It helps inform the response for the 2026 Bundibugyo Virus Disease (BVD) outbreak – from raw patient records to a finished report.
What is Bundibugyo Virus Disease?
Bundibugyo Virus Disease is a rare and severe haemorrhagic fever caused by the Bundibugyo ebolavirus, first identified in Uganda in 2007. Like other members of the Ebola family, it spreads through direct contact with the bodily fluids of an infected person, causes a high case fatality rate, and requires specialist treatment in isolation facilities.
The 2026 outbreak is centred in the Democratic Republic of Congo (DRC). Cases are being reported across multiple health zones, including northeastern provinces of Ituri and Nord-Kivu, with smaller numbers also reported in Sud-Kivu and neighbouring Uganda. The Ministry of Health, Hygiene and Social Welfare of the DRC, and other outbreak response partners, collect surveillance data to track the outbreak.
Why is tracking an outbreak so difficult?
The data we see today is always a blurry picture of the past, not the present. For example, when a person falls ill in a rural area where access to care is limited, days can pass before their case is logged in the surveillance system. By the time a case first appears in the data, the person may have already recovered, died, or passed the disease on to others, and laboratory confirmation of the case often follows later still. Beyond this, some persons may never be tested, which also makes it difficult to get an accurate picture of the state of the outbreaks.
The journey of a single case, as pictured below, shows where the delays and leaks can creep in.
The journey from illness to recorded case
|
🤒 Symptoms appear Patient falls ill |
→ | 📋
Case reported Logged on the linelist |
→ | 🔬
Sample tested At a laboratory |
→ |
✅ Result confirmed Positive or negative |
|
| Onset to report: about 5 days on average | |||||||
Why this matters for the analysis: A key part of our work is building a statistical model of these delays, so that we can assess the lag in the data and estimate what the outbreak looks like today, not just how it looked a week ago when the most recently recorded cases first fell ill. Further problems for the analysis are that some cases are simply never reported, some go untested, and some don’t get a test result at all. It’s important to remember that outbreak data can be incomplete and delayed. Our work aims to carefully turn that imperfect information into evidence that helps response teams plan.
What is a linelist?
All of our analysis starts from a linelist —a table, updated continuously by field teams, that records every known patient in the outbreak. Each row is one person; each column is a piece of information about them.
What a linelist looks like (simplified)
| Patient ID | Date symptoms began | Date case reported | Test result | Health zone |
| 0001 | 2026-05-03 | 2026-05-08 | Positive | Zone A |
| 0002 | 2026-05-05 | 2026-05-09 | Positive | Zone B |
| 0003 | 2026-05-07 | 2026-05-11 | Negative | Zone A |
| ⋮ | ⋮ | ⋮ | ⋮ | ⋮ |
Each row is one patient. Real linelists contain many more columns, but these are the core columns for our analysis.
While the linelist data we receive from partners in DRC is pseudo-anonymised (i.e. no patient names or addresses are included), the sensitivity of the data means we encrypt it before it is shared with the analysis team. – The analysis workflow helps streamline the process of decrypting the data, checking for errors or inconsistencies in the dates, and standardising the column names ready for analysis.
What is the Imperial team building?
The analyses presented here stem from a long-standing scientific collaboration between Imperial College London’s Centre for Global Infectious Disease Analysis and the Institut National de Recherche Biomédicale (INRB). Through this partnership, a bespoke analytical workflow is being developed and implemented to provide timely evidence for the 2026 Bundibugyo Virus Disease outbreak response.
Key contributions of Imperial and its partners
- Design and build the statistical models that estimate how long delays between symptom onset and case recording are, and how they differ across geographic areas of the DRC.
- Build the exponential growth models that estimate how fast the outbreak is growing, and whether the growth rate has changed over time.
- Develop the disease progression projection that translates case growth projections into hospital bed demand, drawing on established clinical assumptions about BVD, where available.
- Build a streamlined analysis workflow that supports regular updates as new outbreak data arrive, helping to ensure results are generated consistently and transparently.
- Implement end-to-end data encryption so that confidential patient records never leave the secure environment unprotected.
What each part of the pipeline does
Step 1: Data in
📥 Preparing the data
The latest outbreak data are received in an encrypted format. The records are securely prepared for analysis by checking for missing or inconsistent information and converting them into a standardised format. Every change is logged, providing a clear record of how the data were processed.
Step 2: Delays
⏳ Accounting for delays in the data
There is often a gap between when someone develops symptoms and when their case appears in the surveillance data. The analysis estimates these reporting delays and accounts for differences between case types and geographic areas, helping to build a more accurate picture of the outbreak.
Step 3: Assessing current trends
📈 Estimating how the outbreak is changing
The analysis estimates how quickly the outbreak is growing or declining over time. It also calculates the reproduction number (R), which indicates whether transmission is increasing, remaining stable or decreasing (see below).
Step 4: Looking ahead
📋 Projecting future healthcare needs
Using current trends, the analysis projects how the outbreak could evolve over the coming weeks. These projections can help inform planning for healthcare resources, including treatment capacity in different geographic areas.
Step 5: Sharing the results
📄 Producing reports for decision-makers
The analysis is compiled into technical reports and presentation materials that summarise the latest findings, including maps, epidemic curves, estimates of transmission and future projections. These outputs are reviewed by analysts before being shared with response partners.
Understanding the reproduction number (R)
The reproduction number, or R, is one of the most important numbers in outbreak analysis. It tells us, on average, how many people each infected person passes the disease to. The workflow estimates R continuously as new data arrives.
|
🔴 R > 1 OUTBREAK GROWING Each person infects more than one other on average. The outbreak will expand unless action is taken. |
🟡
R = 1 HOLDING STEADY Each person infects exactly one other on average. Case numbers remain roughly stable. |
🟢 R < 1 OUTBREAK DECLINING Each person infects fewer than one other on average. Case numbers will fall towards zero. |
The goal of the outbreak response – through case finding, isolation, contact tracing, and community engagement – is to push R below 1 and keep it there. Our workflow tracks whether that is happening.
Note that R and the growth rate measure related but slightly different things: the growth rate is the day-by-day change in case numbers, while R focuses on the chain of transmission between individuals. The workflow estimates both.
What decision makers can do with these results
- Inform healthcare planning: Projections of future case numbers and healthcare needs can help response partners understand where treatment capacity may come under greatest pressure. By highlighting the gap between expected demand and existing capacity, the analysis provides evidence to support operational planning and resource allocation as the situation evolves.
- Highlight uncertainty in the surveillance data: Reporting delays and incomplete data mean that recent case numbers are often uncertain, particularly in areas where few cases have been reported. By estimating these uncertainties, the analysis helps decision maker interpret the available data with appropriate caution and identify where additional information may be most valuable.
- Monitor how the outbreak is changing: By estimating the outbreak’s growth rate and reproduction number (R), the analysis provides an up-to-date picture of transmission over time. These indicators help response teams understand whether the outbreak appears to be growing, stabilising or declining, whilst recognising that many factors influence these trends.
Review updated analyses as new data becomes available: As additional surveillance data is received, the analysis is updated by the teams to provide revised estimates and projections. This enables response partners to base discussions and planning on the latest available evidence while recognising that the analysis is continually refined as new information emerges.
Real-World example
Recently, the pipeline was used to project the expected demand for hospital and isolation beds, under different assumptions of length of stay for suspected patients that then go on to test negative. The results were used to directly inform operational discussions around the optimisation of available BVD diagnostics and isolation beds.
Working together across the outbreak response: The analyses presented here are part of a much broader collaborative effort. Reliable modelling depends on the work of surveillance teams collecting data in the field, laboratory staff processing samples, clinicians caring for patients, data managers maintaining high-quality records, and public health partners interpreting the results alongside local knowledge and operational realities. The greatest value comes from combining these different sources of expertise so that evidence from the analysis can be used alongside on-the-ground experience to support informed outbreak response decisions.
Protecting patient data
All patient records are encrypted throughout the analysis – only team members with a specific cryptographic key can access the underlying data. The results we share with partners and decision makers contain only summary statistics and aggregated estimates, never individual patient information.
