- New feature:
link_encounters()gains an optionalfallback_visit_colargument. By default (NULL), a row missingvisit_colis never merged with another row solely because they share that same missing value (see the NA-key bug fix below) -- correct in general, but it means a real split episode (an ED row and a direct-admit row for the same visit) withVisit_IDmissing on both sides is left as two unmatched rows instead of one linked episode, sincevisit_colalone can't confirm they match.fallback_visit_colnames a secondary identifier (e.g.C_BioSense_ID, observed in real production data to be assigned identically to both rows of such a split episode even whenVisit_IDis missing on both) to use for that specific case: a row missingvisit_colis matched to another row sharingfacility_coland the samefallback_visit_colvalue instead of being left unmatchable. Rows with a realvisit_colvalue are never affected, and omitting the argument preserves the original NA-key behavior exactly. Reported against real production data in a downstream ETL pipeline (149 real split episodes sharing a missingVisit_IDand a matchingC_BioSense_ID, confirmed viadistinct()after dropping the differingHasBeen_-derived field). - Bug fix:
classify_duplicates()no longer emits a spurious base R warning ("replacement element 1 has 1 row to replace 0 rows") on a genuinely clean pull with zero duplicates.janitor::adorn_pct_formatting()errors on a 0-row tabyl;$overallis now built directly as an empty tibble in that case instead of being routed through it. - Bug fix:
dedupe(),summarize_duplicates(),classify_duplicates(), andlink_encounters()no longer collapse rows that share a missingfacility_colorvisit_colvalue into a single group.dplyr::group_by()(and.by =) follow SQL'sGROUP BYconvention of treating everyNAas equal to every otherNAfor grouping purposes, even thoughNA == NAevaluates toNAeverywhere else in R. A missing identifier means a row's true identity is unknown, not confirmed to match every other row with a missing identifier. Previously, several rows sharing a missingVisit_IDat the same facility were silently treated as one duplicated visit:dedupe()discarded all but one of them,summarize_duplicates()/classify_duplicates()reported them as duplicated when they weren't, andlink_encounters()ran its episode-reconciliation logic (has_been_flagmax(), field merging) across genuinely unrelated visits sharing one synthesized.episode_id. Each of these functions now treats a row with a missing key as its own distinct record and emits an informational message (rlang::inform(), suppressible viaverbose = FALSEwhere that argument exists) reporting how many rows were affected. - Behavior change:
dedupe(),summarize_duplicates(),classify_duplicates(),review_facility_ed_visits(), andlink_encounters()now preferHospital/C_BioSense_Facility_IDoverHospitalNameas the defaultfacility_col, wheneverHospitalis present in the data, falling back toHospitalNameonly if it isn't.Hospitalis a stable numeric identifier;HospitalNameis a display string that changes on a facility rename or rebrand, so grouping by name can silently split one facility's rows into two across a rename, or merge two different facilities that briefly share a display name. An explicitly suppliedfacility_colalways overrides this preference exactly as given.filter_care_setting()'sfacility_colis deliberately unchanged, since it matches againstfix_facility_type_vector's exact facility name strings; that function already exposes a separate, ID-preferringfacility_id_col/fix_facility_id_vectorfor the same durability benefit. If you have code that assumesdedupe()(or the other affected functions) group byHospitalName/hospital_nameby default, and your data includesHospital, update it to referencehospital/Hospitalinstead, or passfacility_col = HospitalNameexplicitly to keep the old behavior. The five affected functions'facility_colargument now defaults toNULL(was a fixed column name) so thatargs()/the Usage line accurately reflect that the real default is resolved at runtime rather than printing a fixed default that's no longer accurate; this matches howorder_by/date_colalready behave elsewhere in the package. - Bug fix in
link_encounters(): when derivingpatient_classfromHasBeen_flags (the fallback path used whenC_Patient_Class_Listis absent), thehas_been_e/has_been_admitted/etc. columns were silently lost from the output for any row that never came frominpatient_admission_datadirectly, showing asNAon single-row episodes, and, worse, as an incorrectly reconciled value (e.g.has_been_e = 0on a merged episode that genuinely included an ED visit) on multi-row merged episodes.link_encounters()now preserves each row's true originalHasBeen_values across the pivot, so every episode shows correct0/1values, neverNA, and never an incorrect reconciled value. - Bug fix in
link_encounters(): in the sameHasBeen_-flag fallback path, whened_datacontained bothHasBeenAdmittedandHasBeenI,HasBeenIwas dropped fromed_dataentirely to keep it from contributing a redundant "Admitted" row to the patient-class pivot, which discarded its real0/1values for everyed_datarow. Afterbind_rows()withinpatient_admission_data, only rows sourced from the inpatient pull retained a realhas_been_ivalue; every ED-pull row showedNA.HasBeenIis now excluded from the pivot without being removed from the data, so its true value is preserved on every row. - The message issued when
HasBeenO = 1visits are present is now an informational message (suppressible viaverbose = FALSE) rather than a warning, and no longer implies these visits will showpatient_class = "Outpatient"in the default collapsed output; they won't, whenever the same episode also includes an ED or inpatient-admission record (the norm forlink_encounters()'s two-pull input), since only the primary row'spatient_classsurvives collapsing.HasBeenO = 1is ordinary co-occurring ESSENCE data, not a data quality concern. - Added
CITATION.cff(Citation File Format) at the package root for GitHub's "Cite this repository" feature and Zenodo DOI metadata, alongside the existinginst/CITATIONused bycitation("sysPrep"). - Functions validated against the NSSP ESSENCE
va_er(Patient Location, Full Details) andva_hosp(Facility Location, Full Details) data sources. dedupe(): Remove duplicate ESSENCE records with flexible keep strategy.summarize_duplicates(): Summarize duplicate counts by facility.classify_duplicates(): Classify duplication mechanism by type. Supportsverboseto suppress informational messages.filter_care_setting(): Filter to valid emergency and inpatient care settings. Supportsverboseto suppress informational messages, andfix_facility_id_vectorto correct known facilities by their stableHospital/C_BioSense_Facility_IDvalue, more durable across facility name changes thanfix_facility_type_vector.link_encounters(): Link ED and inpatient encounters into care episodes, merging each episode's rows into one composite row by default (return_format = "collapsed");HasBeen_flags reconciled via max, andCCDD/CCDDParsed/CCDDCategory_flat/C_Death/Discharge_Disposition/DispositionCategoryreconciled via configurablemerge_fieldsstrategies.return_format = "long"preserves the prior unmerged output. Supportsverboseto suppress informational messages. Breaking:inpatient_admission_datais now required; the prior single-pull mode (ed_dataalone) could not detect a genuine direct admission (structurally absent from aHasBeenE = 1pull) and was a no-op on an already-deduplicated ED-to-inpatient escalation, solink_encounters()now aborts with an actionable message instead of silently returninged_dataunchanged. Query a second ESSENCE pull filtered toHasBeenAdmitted = 1(orHasBeenI = 1), deduplicate it separately, and pass it asinpatient_admission_data.- Added
essence_ed_rawandessence_inp_raw: two small synthetic datasets representing separately queriedHasBeenE = 1andHasBeenAdmitted = 1ESSENCE pulls, used byvignette("encounter-linkage")andlink_encounters()'s own examples to demonstrate two-pull linkage.essence_raw/essence_cleanare unchanged by this and continue to represent a single realistic ED pull. review_facility_ed_visits(): Flag facility visit count outliers for QA. Supportsverboseto suppress informational messages.assign_treating_geography(): Assign treating facility geography to out-of-state visits. By default writes to newnew_region_col/new_zip_colcolumns ("region_hybrid"/"zip_code_hybrid"), leavingregion_col/zip_coluntouched; setoverwrite = TRUEto overwrite them in place instead. Supportsverboseto suppress informational messages.assign_facility_geography(): Assign facility geography to all visits. Samenew_region_col/new_zip_col/overwritebehavior asassign_treating_geography(), defaulting to"region_facility"/"zip_code_facility". Supportsverboseto suppress informational messages.