Data Freshness & Availability
Understanding when data becomes available in iPCV is important when designing reports, dashboards, analytical pipelines and data extraction processes.
This page explains how data moves from source systems into iPCV, when data becomes available, and the factors that can influence data freshness.
How Data Becomes Available in iPCV
Section titled “How Data Becomes Available in iPCV”iPCV data originates from EMIS Web and passes through several processing stages before it is available for reporting and analytics:
- EMIS Web
- Replication & Ingestion
- EXA Data Lake
- iPCV Transformation & Validation
- iPCV Data Models
- Customer Reporting & Analytics
As data passes through multiple processing stages, information is not immediately available within iPCV after it is recorded in the source clinical system.
Data Availability Timeline
Section titled “Data Availability Timeline”Processing runs daily in iPCV. The table below provides indicative processing timelines for data entering iPCV through the required processing layers.
| Stage | Typical Timeline |
|---|---|
| Data entered in EMIS Web | Day 0 |
| Replicated into EXA Data Lake | Typically within 4–6 hours |
| Processed through iPCV ETL pipelines | Daily processing cycle, triggered overnight |
| Available within iPCV models | Typically within ~16 hours of source system activity |
These timelines are provided as guidance and may vary depending on source system activity, replication schedules and operational processing windows. The timing of the overnight processing cycle is chosen to minimise data latency for the previous day. If a record is not available in the datalake when processing starts, it will be picked up in the following run. Records entered in the evening and overnight are therefore more likely to have higher latency than records entered during the working day.
It is important to understand the dates used in the clinical record and iPCV processing. A record appearing today can carry an event date from months or years in the past, because that date is entered by the clinician and reflects when the event happened, not when the entry was created. See Dates and Time Fields for the full explanation.
What This Means In Practice
Section titled “What This Means In Practice”As a general rule:
- Data entered into EMIS Web during the day should be expected to appear in iPCV after the next processing cycle has completed.
- Recently entered information will not be available immediately.
- Customers should treat iPCV as a regularly refreshed analytical dataset rather than a real-time operational system.
When designing operational reports or analytics, it is important to account for this expected latency.
Scheduling Data Extractions
Section titled “Scheduling Data Extractions”When building automated reporting or extraction processes, schedule data loads around expected iPCV refresh cycles. Data refreshes for individual iPCV data models are carried out in parallel. The data is made available to customers as soon as each model is completed, so customers do not need to wait for all models to complete before starting their data extractions.
Each iPCV flavour has a table called data_freshness that customers can use to
detect when a model is ready for extraction.
Data Consistency Between Models
Section titled “Data Consistency Between Models”Data is continuously ingested from source systems into the EXA datalake and the iPCV pipeline picks up the data that is available when the daily processing cycle starts. The refreshed data in the iPCV models becomes available as soon as that model has completed each day.
This means that a record with a foreign key can become available to customers before the corresponding primary key row is available in another model. The primary key row may arrive when the referenced model completes its daily refresh, or it may arrive with the following day’s refresh, depending on whether both records were available in the datalake when the iPCV pipeline commenced.
Customers querying the data will need to take this into account when loading data. The data has eventual consistency but consistency cannot be guaranteed for newly arriving data.
If sensitive or confidential record filtering is applied in a customer’s configuration, this can affect the visibility of linked records. See Privacy & Filtering Controls for general information about patient and record filtering. It is therefore possible that a foreign key is present that references a primary key record which is hidden — for example, a consultation section could be visible with a foreign key that cannot be resolved because the referenced observation is hidden.
Why Record Counts Can Change
Section titled “Why Record Counts Can Change”It is normal for record counts to change between reporting periods. Changes may occur due to:
- New clinical activity being recorded
- Updates or corrections to existing records
- Registration status changes
- Patient consent changes
- Privacy and filtering controls being applied
- Previously unavailable data arriving in a later processing cycle
When validating differences between reporting runs, both data freshness and applicable filtering rules should be considered. See Privacy & Filtering Controls for more information on filtering.
Frequently Asked Questions
Section titled “Frequently Asked Questions”Why can’t I see information that was entered today?
Section titled “Why can’t I see information that was entered today?”It hasn’t reached iPCV yet. Data is replicated and then processed through iPCV on a daily cycle, so recently entered information may not appear until the next cycle completes.
Why did my record counts change since yesterday?
Section titled “Why did my record counts change since yesterday?”Record counts can change as new information becomes available, source records are updated, patient status changes occur, or privacy and consent rules affect which records are included in your dataset.
Can iPCV be used for real-time operational reporting?
Section titled “Can iPCV be used for real-time operational reporting?”No. iPCV is designed as a daily refreshed analytical dataset and is updated through scheduled processing cycles rather than real-time transactions. Customers should account for expected latency when building reports and dashboards.
Why can I not see a record that is referenced by a foreign key?
Section titled “Why can I not see a record that is referenced by a foreign key?”See Data Consistency Between Models above.
What should I do if expected data is missing?
Section titled “What should I do if expected data is missing?”Before raising an investigation, consider:
- Whether the expected refresh cycle has completed.
- Whether the source data was entered before the processing window.
- Whether any privacy or filtering controls may affect the records returned.
- Whether the data is being queried from the correct model.
How do I know when a model was last updated?
Section titled “How do I know when a model was last updated?”Use the data_freshness table available in each iPCV flavour. This table shows
when a data model was last updated and can be used to check when a model is
ready for extraction. Delta tables are not listed separately in data_freshness
as they become available at the same time as the main tables.