Repository navigation
Questions about EnsembleStat line type output when using observation uncertainty #3202
Replies: 1 comment 1 reply
|
Marion, thanks for writing up this discussion. These are great questions about the impact of enabling observation uncertainty in Ensemble-Stat. You note that it's not entirely clear and obvious what Ensemble-Stat outputs are impacted by observational uncertainty. This raises some questions about configuration, metadata, and documentation. First, I'll note that this is not an isolated issue. While the MET output DOES include some limited metadata in the first 24 STAT header columns, it is NOT comprehensive. It's just not feasible to add new ASCII columns to the STAT output lines to describe every possible configuration choice. As you know, some config choices have a big impact on results, including regridding methods, censoring and/or filtering input data/pairs, climatology choice, applying topography filters or corrections. None of these choices appear directly in those header columns. All of these configuration options prompted us to add the user-settable There have been some requests for MET to write STAT output to NetCDF files, which could be more efficient and also provide an opportunity for writing more complete metadata. But it also introduces other challenges. Regarding your recommendations:
I totally agree. We can improve it for the next version.
I don't think this is a good idea since it'd really complicate the logic of parsing, plotting, and loading Ensemble-Stat output into a database. I understand that you want to clearly see the impact of adding observation uncertainty, primarily because you're currently investigating that. Someone else investigating the impact of topography corrections would be interested in results with and without it enabled. For the time being, I'd strongly recommend using the I wholeheartedly agree with the need for improved documentation and the need to investigate the missing and/or inconsistent column names you reported. And I'll write up an issue to that effect (dtcenter/MET#3327). I'm less convinced that we need Ensemble-Stat to produce a full set of both perturbed and unperturbed output in a single run. Just like regridding methods, climatology choices, topography corrections, and many other configuration options, they do warrant careful consideration. But once the an informed decision has been made, it's unlikely you'd need a full set of output for both options from each run. This also highlights the value of METplus Use Cases. Publishing your METplus configuration really supports repeatable results by providing a record of the many configuration option that go into each run. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
I have been inspecting the output from EnsembleStat with observation uncertainty.
I thought that I could get away with not running EnsembleStat without observation uncertainty but unfortunately not all the output (with and without) is consistently written to file when observation uncertainty is switched on. For some line types only the perturbed is written to file, but you wouldn't know this is the perturbed output by just looking at the output.
The documentation doesn't explicitly say that the output such as in rhist and orank etc is for perturbed ensembles only when obs uncertainty is switched on. i confirmed this by running the same data through EnsembleStat twice: with and without. The numbers are different (as one might expect/hope). But if the output is somehow detached from the directory (which you hopefully gave a helpful name), it is hard to know whether the output is with or without obs uncertainty (unless you reran the job). What users do and don't do can be hard to predict! It would help/be useful if the columns (or column names) are amended or appended to reflect this (i.e. both outputs are provided in all line types).
The ecnt line type does enable you to compare the ME to the ME_OERR etc Whilst it is also noteworthy that the IGN_OERR is included, I don't see a CRPS_OERR? This seems very odd to me. Of course, it could be calculated from the ORANK values (albeit with some pain and may be inconsistent with internal calculations). I can't think why this is the case? But there is still the problem that the output of some line types is not self-describing enough. Being sure that you are using the perturbed oranks rather depends on you doing a good job at naming files/directories. There is a related issue that it is also impossible to tell whether the ensemble control has been included or excluded from the stat calculations. I think this has been raised before. This would seem to need another column perhaps with a flag (though this isn't totally desirable as it inflates the file sizes with a constant value but I can't think of another option).
Could you please consider a few core enhancements to the EnsembleStat line types to reduce ambiguity and enhance clarity?
In summary, there seems to be an inconsistency in outputs provided when observation uncertainty is included and an element of ambiguity which ought to be improved to prevent users getting confused when plotting and analysing results.
All reactions