Overview
This section provides the foundational knowledge required to successfully work with iPCV data. Before exploring individual data models, it is important to understand how data is refreshed, how changes are delivered, and the key operational concepts that influence how datasets should be consumed and maintained.
Understanding Explorer
Section titled “Understanding Explorer”Explorer is the SQL endpoint used to connect to the datalake and access models including iPCV. It gives customers access to a powerful analytical engine to run queries on data in-situ and download the relevant results. This limits the unnecessary movement of patient data and allows customers to shape the data as they extract it to better meet their use case. Alternatively, full datasets can also be extracted through Explorer if a customer requires an up-to-date local copy in their own data warehouse.
Understanding Data Availability
Section titled “Understanding Data Availability”iPCV is designed for analytics, reporting and research workloads rather than real-time operational use.
Data from source clinical systems is replicated, processed and transformed before becoming available through the iPCV models. As a result, there will always be a delay between an event being recorded in a source system and that event becoming available within iPCV.
Understanding data refresh behaviour is important when:
- Building reporting solutions
- Designing data pipelines
- Performing trend analysis
- Validating record counts
- Comparing data across different reporting periods
For more information, see Data Freshness & Availability.
Understanding Deltas
Section titled “Understanding Deltas”One of the key capabilities of iPCV is support for incremental processing through delta-based data consumption.
Rather than reloading complete datasets each day, customers can identify and process only records that have changed since their previous extraction. This approach can significantly reduce processing time, storage requirements and query volumes.
Delta processing is commonly used when:
- Maintaining reporting databases
- Building data warehouses
- Synchronising local copies of data
- Running scheduled analytical processes
For more information, see Understanding Deltas.
Maintaining a Local Copy of iPCV Data
Section titled “Maintaining a Local Copy of iPCV Data”Many customers choose to maintain a local copy of iPCV data within their own environments. A typical approach consists of:
- Performing an initial full extraction.
- Storing the data locally.
- Processing regular delta updates.
- Applying changes to maintain an up-to-date local dataset.
This approach allows organisations to build reporting and analytics solutions whilst reducing the need for repeated full data extractions.
Understanding Privacy & Filtering
Section titled “Understanding Privacy & Filtering”Depending on your organisation’s configuration and data sharing agreements, iPCV may include filtering and privacy controls that affect the data available to you.
These controls can include:
- Patient-level filtering
- Consent-based filtering
- Sensitive record exclusions
- Sensitive code exclusions
- Confidentiality-based restrictions
Understanding how these controls affect record counts and data visibility is particularly important when comparing datasets between organisations or validating analytical outputs.
For more information, see Privacy & Filtering Controls.
Building Efficient Queries
Section titled “Building Efficient Queries”iPCV has been designed to simplify analytical querying by presenting commonly used healthcare concepts as analyst-friendly models.
Before developing large analytical workloads, review:
- Query Examples & Common Patterns
- Trino Query Optimisation
- Individual data model documentation in the Models section