diff --git a/docs/Contributors_Guide/code_profiling.rst b/docs/Contributors_Guide/code_profiling.rst index ae85e3bf25..185afecdec 100644 --- a/docs/Contributors_Guide/code_profiling.rst +++ b/docs/Contributors_Guide/code_profiling.rst @@ -2,16 +2,14 @@ Code Profiling ************** -Benchmarking (also referred to as profiling) of MET tools is accomplished using the CTRACK tool: - https://github.com/Compaile/ctrack +Benchmarking (also referred to as profiling) of MET tools is accomplished using the `CTRACK tool `_. -This code is licensed under the MIT License: - https://github.com/Compaile/ctrack/blob/main/LICENSE +This code is licensed under the `MIT License `_. Benchmarking uses a macro and the C++ source code is readily instrumented by including ctrack.hpp and by adding **CTRACK** at the top of the function of interest. By default, the tool generates summary and detail metrics to stdout (standard output) -in easy to read, well-formatted tables. The ctrack.hpp file has been modified to permit +in easy-to-read, well-formatted tables. The ctrack.hpp file has been modified to permit saving these tables to their respective text files (summary_output.txt and detail_output.txt). @@ -41,8 +39,10 @@ Code that is currently instrumented The following code is instrumented using CTRACK: - MET/src/basic/vx_util/main.cpp - - do_pre_process function - - do_post_process function + + - do_pre_process function + - do_post_process function + - MET/src/tools/core/ensemble_stat/ensemble_stat.cc - MET/src/tools/core/ensemble_stat/ensemble_stat_conf.cc - MET/src/tools/other/grid_diag/grid_diag.cc @@ -50,12 +50,12 @@ The following code is instrumented using CTRACK: Benchmarking with Python script ------------------------------- -The benchmarking.py script invokes MET code either via **MET command line commands** or **METplus use cases** as +The benchmark.py script invokes MET code either via **MET command line commands** or **METplus use cases** as specified by the **run_met_directly** setting in the benchmark.yaml configuration file. The metrics from the summary and detail tables are consolidated into csv and tabular text files (the locations of these consolidated metrics text files are specified in the benchmark.yaml configuration file). The CTRACK summary_output.txt and detail_output.txt reports (containing the performance metrics) are written to the directory -from which the benchmarking.py script was executed. An information file is also generated that captures the version +from which the benchmark.py script was executed. An information file is also generated that captures the version of Python used, a timestamp, and any other relevant information for capturing the environment under which the code was profiled/benchmarked. @@ -64,138 +64,138 @@ Overview of Steps for Performing Benchmarking 1. .. dropdown:: Instrument the MET code of interest - .. note:: + .. note:: - The ctrack.hpp file is saved in the $HOME/MET/src/basic/vx_util directory and does not need to be modified - or added to any other location. This version of ctrack.hpp has been modified to write the summary and detail - tables to text files. By default, CTRACK is disabled and is enabled at compilation time via the - :code:`--enable-profiler` flag. + The ctrack.hpp file is saved in the $HOME/MET/src/basic/vx_util directory and does not need to be modified + or added to any other location. This version of ctrack.hpp has been modified to write the summary and detail + tables to text files. By default, CTRACK is disabled and is enabled at compilation time via the + :code:`--enable-profiler` flag. - $HOME refers to the path to where the MET source code is saved. + $HOME refers to the path to where the MET source code is saved. - The ctrack.hpp file must be included in the source code of interest: + The ctrack.hpp file must be included in the source code of interest: - .. code-block:: ini + .. code-block:: ini - #ifdef WITH_PROFILER - #include "ctrack.hpp" - #endif + #ifdef WITH_PROFILER + #include "ctrack.hpp" + #endif - The CTRACK directive is placed at the top of the function of interest. Use the preprocessor directive for WITH_PROFILER: + The CTRACK directive is placed at the top of the function of interest. Use the preprocessor directive for WITH_PROFILER: - e.g. ensemble_stat.cc: + e.g., ensemble_stat.cc: - .. code-block:: ini + .. code-block:: ini - void process_grid(const Grid &fcst_grid) { - #ifdef WITH_PROFILER - CTRACK; - #endif - Grid obs_grid; - ... more code + void process_grid(const Grid &fcst_grid) { + #ifdef WITH_PROFILER + CTRACK; + #endif + Grid obs_grid; + ... more code - and the *ctrack::result_print* is placed within the corresponding MET tool's - **main()/met_main()** function + and the *ctrack::result_print* is placed within the corresponding MET tool's + **main()/met_main()** function - e.g. ensemble_stat.cc + e.g., ensemble_stat.cc - .. code-block:: ini + .. code-block:: ini - int met_main(int argc, char *argv[]) { + int met_main(int argc, char *argv[]) { - // Process the command line arguments - process_command_line(argc, argv); + // Process the command line arguments + process_command_line(argc, argv); - // Check for valid ensemble data - process_n_vld(); + // Check for valid ensemble data + process_n_vld(); - // Perform verification - process_vx(); + // Perform verification + process_vx(); - // Save the CTRACK metrics - #ifdef WITH_PROFILER - ctrack::result_print(); - #endif + // Save the CTRACK metrics + #ifdef WITH_PROFILER + ctrack::result_print(); + #endif - .. note :: + .. note :: - The summary_output.txt and detail_output.txt files will only be saved when the ctrack::result_print() function is - called within main() or met_main(). + The summary_output.txt and detail_output.txt files will only be saved when the ctrack::result_print() function is + called within main() or met_main(). 2. .. dropdown:: Compile MET code - .. dropdown:: Configure + .. dropdown:: Configure - From the $HOME/MET directory: + From the $HOME/MET directory: - * source ./internal/scripts/environment/development.xyz - * *xyz* is the name of the host + * source ./internal/scripts/environment/development.xyz + * *xyz* is the name of the host - **By default, CTRACK is disabled**. Enable it with the --enable-profiler option. + **By default, CTRACK is disabled**. Enable it with the --enable-profiler option. - Run one of the following configure commands (to enable all the components and the CTRACK macro): + Run one of the following configure commands (to enable all the components and the CTRACK macro): - .. code-block:: ini + .. code-block:: ini - ./configure --prefix=`pwd` --enable-grib2 --enable-modis --enable-lidar2nc --enable-python --enable-ugrid --enable-profiler + ./configure --prefix=`pwd` --enable-grib2 --enable-modis --enable-lidar2nc --enable-python --enable-ugrid --enable-profiler - or + or - .. code-block:: ini + .. code-block:: ini - ./configure --prefix=`pwd` --enable-all --enable-ugrid --enable-profiler + ./configure --prefix=`pwd` --enable-all --enable-ugrid --enable-profiler - .. dropdown:: Run make install and test + .. dropdown:: Run make install and test - Redirect the output to a log file named make.log: + Redirect the output to a log file named make.log: - .. code-block:: ini + .. code-block:: ini - make install test >& make.log & - tail -f make.log + make install test >& make.log & + tail -f make.log - .. dropdown:: Verify that the expected code is being measured + .. dropdown:: Verify that the expected code is being measured - The summary and detail tables are generated during the MET build (when running the test target). - These tables created by CTRACK can be viewed in the make.log before they are consolidated. - Use the *cat* (concatenation) tool to view the make.log file to view the CTRACK-generated metrics tables that - correspond to the MET tool that was instrumented. + The summary and detail tables are generated during the MET build (when running the test target). + These tables created by CTRACK can be viewed in the make.log before they are consolidated. + Use the *cat* (concatenation) tool to view the make.log file to view the CTRACK-generated metrics tables that + correspond to the MET tool that was instrumented. - .. note:: + .. note:: The CTRACK output is formatted using *BeautifulTable*. Therefore **concatenation** (vs viewing via a text editor like vim) facilitates viewing the human-readable version of the tables. The human-readable form of the tables is also available while running the *tail -f* command when viewing the make.log during compilation. - From the command line: + From the command line: - .. code-block:: ini + .. code-block:: ini - cat make.log + cat make.log - CTRACK summary and detail tables will appear in the make.log file. A - summary table will look like the following: + CTRACK summary and detail tables will appear in the make.log file. A + summary table will look like the following: - .. code-block:: ini + .. code-block:: ini - Summary - +---------------------+---------------------+------------+---------------+-----------------+ - | Start | End | time total | time ctracked | time ctracked % | - +---------------------+---------------------+------------+---------------+-----------------+ - | 2025-04-22 22:43:22 | 2025-04-22 22:44:31 | 69.43 s | 353.42 mcs | 0.00% | - +---------------------+---------------------+------------+---------------+-----------------+ - +----------+-----------------+------+-------+-----------+------------+----------------+---------------+ - | filename | function | line | calls | ae[1-99]% | ae[0-100]% | time ae[0-100] | time a[0-100] | - +----------+-----------------+------+-------+-----------+------------+----------------+---------------+ - | main.cc | do_pre_process | 97 | 1 | 0.00% | 0.00% | 322.08 mcs | 322.08 mcs | - +----------+-----------------+------+-------+-----------+------------+----------------+---------------+ - | main.cc | do_post_process | 119 | 1 | 0.00% | 0.00% | 31.34 mcs | 31.34 mcs | - +----------+-----------------+------+-------+-----------+------------+----------------+---------------+ + Summary + +---------------------+---------------------+------------+---------------+-----------------+ + | Start | End | time total | time ctracked | time ctracked % | + +---------------------+---------------------+------------+---------------+-----------------+ + | 2025-04-22 22:43:22 | 2025-04-22 22:44:31 | 69.43 s | 353.42 mcs | 0.00% | + +---------------------+---------------------+------------+---------------+-----------------+ + +----------+-----------------+------+-------+-----------+------------+----------------+---------------+ + | filename | function | line | calls | ae[1-99]% | ae[0-100]% | time ae[0-100] | time a[0-100] | + +----------+-----------------+------+-------+-----------+------------+----------------+---------------+ + | main.cc | do_pre_process | 97 | 1 | 0.00% | 0.00% | 322.08 mcs | 322.08 mcs | + +----------+-----------------+------+-------+-----------+------------+----------------+---------------+ + | main.cc | do_post_process | 119 | 1 | 0.00% | 0.00% | 31.34 mcs | 31.34 mcs | + +----------+-----------------+------+-------+-----------+------------+----------------+---------------+ @@ -205,8 +205,8 @@ Overview of Steps for Performing Benchmarking .. note:: - the benchmark.py and benchmark.yaml files **must** reside in the same directory - (the benchmark.yaml file does **NOT** need to be specified at the command line) + The benchmark.py and benchmark.yaml files **must** reside in the same directory + (the benchmark.yaml file does **NOT** need to be specified at the command line) .. dropdown:: The following is an example benchmark.yaml config file that utilizes environment variables and full directory paths @@ -218,7 +218,7 @@ Overview of Steps for Performing Benchmarking # # filename - # Timestamp in ISO 1806 format is used to generate output filename + # Timestamp in ISO 8601 format is used to generate output filename # If filename setting is empty string, then timestamp is used. # Otherwise, the specified filename followed by the timestamp will # be used for the output filename. @@ -271,105 +271,105 @@ Overview of Steps for Performing Benchmarking .. dropdown:: Config settings for running via MET command: - - benchmark_output_path + - benchmark_output_path - - **required** - - output directory where the output files will be saved - - specify in one of two ways: + - **required** + - output directory where the output files will be saved + - specify in one of two ways: - - setting the BENCHMARK_OUTPUT_BASE env variable - - explicitly setting the **full** directory path + - setting the BENCHMARK_OUTPUT_BASE env variable + - explicitly setting the **full** directory path - - filename + - filename - - **optional** - - the supplied filename prepended with a Timestamp that follows ISO 1806 format - - if left empty, the timestamp alone will be used as the filename + - **optional** + - the supplied filename followed by a timestamp that follows ISO 8601 format + - if left empty, the timestamp alone will be used as the filename - - run_met_directly + - run_met_directly - - **required** - - set to **True** + - **required** + - set to **True** - - met_command + - met_command - - **required** - - the command to run the MET tool with the appropriate arguments - - this is the same command that would be ordinarily used when running a - MET tool from the command line - - make sure the specified *-outdir* directory exists + - **required** + - the command to run the MET tool with the appropriate arguments + - this is the same command that would be ordinarily used when running a + MET tool from the command line + - make sure the specified *-outdir* directory exists - - met_subdir_name + - met_subdir_name - - **optional** - - if left empty, the consolidated benchmark metrics will be saved to a subdirectory (in the - benchmark_output_path) named after the MET tool + - **optional** + - if left empty, the consolidated benchmark metrics will be saved to a subdirectory (in the + benchmark_output_path) named after the MET tool - - num_runs + - num_runs - - **optional** - - to be used for stress-testing/running command multiple times - - if not set, default value is 1 + - **optional** + - to be used for stress-testing/running command multiple times + - if not set, default value is 1 .. dropdown:: Config settings for running via METplus usecase(s): - - benchmark_output_path + - benchmark_output_path - - **required** - - output directory where the output files will be saved - - specify in one of two ways: + - **required** + - output directory where the output files will be saved + - specify in one of two ways: - - setting the BENCHMARK_OUTPUT_BASE env variable - - explicitly setting the full directory path + - setting the BENCHMARK_OUTPUT_BASE env variable + - explicitly setting the full directory path - - filename + - filename - - **optional** - - the supplied filename prepended with a Timestamp that follows ISO 1806 format - - if left empty, the timestamp alone will be used as the filename + - **optional** + - the supplied filename followed by a timestamp that follows ISO 8601 format + - if left empty, the timestamp alone will be used as the filename - - run_met_directly + - run_met_directly - - **required** - - set to **False** + - **required** + - set to **False** - - metplus_base + - metplus_base - - **required** - - location of the METplus source code, specified by one of the following methods: + - **required** + - location of the METplus source code, specified by one of the following methods: - - indicated as a full path e.g. /home/username/METplus - - setting the METPLUS_BASE environment variable and use the current environment syntax like the following: + - indicated as a full path e.g., /home/username/METplus + - setting the METPLUS_BASE environment variable and using the current environment syntax like the following: - .. code-block:: ini + .. code-block:: ini - !ENV '${SOME_ENV_NAME}' + !ENV '${SOME_ENV_NAME}' - Make sure that the *SOME_ENV_NAME* environment variable is defined + Make sure that the *SOME_ENV_NAME* environment variable is defined - - system.conf + - system.conf - - **required** - - file location of the system.conf file - - full path and file name - - pre-condition: generate a valid system.conf file + - **required** + - file location of the system.conf file + - full path and file name + - pre-condition: generate a valid system.conf file - - wrapper_conf + - wrapper_conf - - **required** - - the location of the METplus wrapper use case config file(s) - - more than one use case can be run - - full path and file name - - pre-condition: generate the necessary wrapper config file(s) + - **required** + - the location of the METplus wrapper use case config file(s) + - more than one use case can be run + - full path and file name + - pre-condition: generate the necessary wrapper config file(s) - - num_runs + - num_runs - - **not yet supported** - - to be used for stress-testing/running command multiple times - - set to 1 + - **not yet supported** + - to be used for stress-testing/running command multiple times + - set to 1 - .. note:: + .. note:: A subdirectory under the output base directory (specified in benchmark_output_path) is created for each use case (based on the use case config filename). @@ -378,88 +378,88 @@ Overview of Steps for Performing Benchmarking 4. .. dropdown:: Invoke the Python script *benchmark.py* to collect the benchmarking metrics - .. note:: + .. note:: Use Python 3.12 or above for running the benchmark.py script - **Pre-conditions:** + **Pre-conditions:** - .. dropdown:: Running MET command + .. dropdown:: Running MET command - Define any necessary environment variables for the corresponding MET tool (e.g. Ensemble-Stat tool environment + Define any necessary environment variables for the corresponding MET tool (e.g., Ensemble-Stat tool environment variables specified in the $HOME/METplus/metplus/parm/met_config/EnsembleStatConfig_wrapped) - .. dropdown:: Example Ensemble-Stat config + .. dropdown:: Example Ensemble-Stat config - .. code-block:: ini + .. code-block:: ini - #!/usr/bin/bash - - export METPLUS_CENSOR_THRESH=""; - export METPLUS_CENSOR_VAL=""; - export METPLUS_CI_ALPHA="ci_alpha = [0.05];"; - export METPLUS_CLIMO_CDF_DICT=""; - export METPLUS_CLIMO_MEAN_DICT=“”; - export METPLUS_CLIMO_STDEV_DICT=""; - export METPLUS_CONTROL_ID=""; - export METPLUS_DESC="desc = \"NA\";"; - export METPLUS_DUPLICATE_FLAG=""; - export METPLUS_ECLV_POINTS=""; - export METPLUS_ENS_MEMBER_IDS=""; - export METPLUS_ENS_PHIST_BIN_SIZE=""; - export METPLUS_ENS_SSVAR_BIN_SIZE=""; - export METPLUS_ENS_THRESH="ens_thresh = 1.0;"; - export METPLUS_FCST_CLIMO_STDEV_DICT=""; - export METPLUS_FCST_FIELD="field = [{ name=\"APCP\"; level=\"A01\"; }];"; - export METPLUS_FCST_FILE_TYPE=""; export METPLUS_GRID_WEIGHT_FLAG=""; - export METPLUS_INTERP_DICT="interp = {vld_thresh = 1.0;shape = SQUARE;type = {method = [NEAREST];width = [1];}}"; - export METPLUS_MASK_GRID=""; - export METPLUS_MASK_POLY=""; - export METPLUS_MESSAGE_TYPE=""; - export METPLUS_MET_CONFIG_OVERRIDES=""; - export METPLUS_MODEL="model = \"RRFS\";"; - export METPLUS_NC_ORANK_FLAG_DICT="nc_orank_flag = {latlon = TRUE;mean = TRUE;raw = TRUE;rank = TRUE;pit = TRUE;vld_count = TRUE;weight = FALSE;}"; - export METPLUS_OBS_CLIMO_MEAN_DICT=""; - export METPLUS_OBS_CLIMO_STDEV_DICT=""; - export METPLUS_OBS_ERROR_FLAG=""; - export METPLUS_OBS_FIELD="field = [{ name=\"APCP\"; level=\"A01\"; }];"; - export METPLUS_OBS_FILE_TYPE=""; export METPLUS_OBS_QUALITY_EXC=""; - export METPLUS_OBS_QUALITY_INC=""; export METPLUS_OBS_THRESH=""; - export METPLUS_OBS_WINDOW_DICT="obs_window = {beg = -1800;end = 1800;}"; - export METPLUS_OBTYPE="obtype = \"CCPA\";"; - export METPLUS_OBTYPE_AS_GROUP_VAL_FLAG=""; - export METPLUS_OUTPUT_FLAG_DICT="output_flag = {ecnt = NONE;rps = NONE;rhist = STAT;phist = STAT;orank = STAT;ssvar = STAT;relp = STAT;}"; - export METPLUS_OUTPUT_PREFIX=""; - export METPLUS_POINT_WEIGHT_FLAG=""; - export METPLUS_PROB_CAT_THRESH=""; - export METPLUS_PROB_PCT_THRESH=""; - export METPLUS_REGRID_DICT="regrid = {to_grid = OBS;method = NEAREST;width = 1;vld_thresh = 0.5;shape = SQUARE;}"; - export METPLUS_SKIP_CONST=""; exp - - - .. dropdown:: Running via METplus Usecase(s) - - Define the necessary environment variables that are required for running any METplus use case. - - - **Running the Python script** - - Run the following from the command line (from the location where the benchmark.py file is located): + #!/usr/bin/bash + + export METPLUS_CENSOR_THRESH=""; + export METPLUS_CENSOR_VAL=""; + export METPLUS_CI_ALPHA="ci_alpha = [0.05];"; + export METPLUS_CLIMO_CDF_DICT=""; + export METPLUS_CLIMO_MEAN_DICT=“”; + export METPLUS_CLIMO_STDEV_DICT=""; + export METPLUS_CONTROL_ID=""; + export METPLUS_DESC="desc = \"NA\";"; + export METPLUS_DUPLICATE_FLAG=""; + export METPLUS_ECLV_POINTS=""; + export METPLUS_ENS_MEMBER_IDS=""; + export METPLUS_ENS_PHIST_BIN_SIZE=""; + export METPLUS_ENS_SSVAR_BIN_SIZE=""; + export METPLUS_ENS_THRESH="ens_thresh = 1.0;"; + export METPLUS_FCST_CLIMO_STDEV_DICT=""; + export METPLUS_FCST_FIELD="field = [{ name=\"APCP\"; level=\"A01\"; }];"; + export METPLUS_FCST_FILE_TYPE=""; export METPLUS_GRID_WEIGHT_FLAG=""; + export METPLUS_INTERP_DICT="interp = {vld_thresh = 1.0;shape = SQUARE;type = {method = [NEAREST];width = [1];}}"; + export METPLUS_MASK_GRID=""; + export METPLUS_MASK_POLY=""; + export METPLUS_MESSAGE_TYPE=""; + export METPLUS_MET_CONFIG_OVERRIDES=""; + export METPLUS_MODEL="model = \"RRFS\";"; + export METPLUS_NC_ORANK_FLAG_DICT="nc_orank_flag = {latlon = TRUE;mean = TRUE;raw = TRUE;rank = TRUE;pit = TRUE;vld_count = TRUE;weight = FALSE;}"; + export METPLUS_OBS_CLIMO_MEAN_DICT=""; + export METPLUS_OBS_CLIMO_STDEV_DICT=""; + export METPLUS_OBS_ERROR_FLAG=""; + export METPLUS_OBS_FIELD="field = [{ name=\"APCP\"; level=\"A01\"; }];"; + export METPLUS_OBS_FILE_TYPE=""; export METPLUS_OBS_QUALITY_EXC=""; + export METPLUS_OBS_QUALITY_INC=""; export METPLUS_OBS_THRESH=""; + export METPLUS_OBS_WINDOW_DICT="obs_window = {beg = -1800;end = 1800;}"; + export METPLUS_OBTYPE="obtype = \"CCPA\";"; + export METPLUS_OBTYPE_AS_GROUP_VAL_FLAG=""; + export METPLUS_OUTPUT_FLAG_DICT="output_flag = {ecnt = NONE;rps = NONE;rhist = STAT;phist = STAT;orank = STAT;ssvar = STAT;relp = STAT;}"; + export METPLUS_OUTPUT_PREFIX=""; + export METPLUS_POINT_WEIGHT_FLAG=""; + export METPLUS_PROB_CAT_THRESH=""; + export METPLUS_PROB_PCT_THRESH=""; + export METPLUS_REGRID_DICT="regrid = {to_grid = OBS;method = NEAREST;width = 1;vld_thresh = 0.5;shape = SQUARE;}"; + export METPLUS_SKIP_CONST=""; exp + + + .. dropdown:: Running via METplus Usecase(s) + + Define the necessary environment variables that are required for running any METplus use case. + + + **Running the Python script** + + Run the following from the command line (from the location where the benchmark.py file is located): - .. note:: + .. note:: An AssertionError message is printed to the terminal if the benchmark.py script is not run in the $BASE/MET/internal/scripts/benchmark directory. - .. code-block:: ini + .. code-block:: ini - cd $BASE/MET/internal/scripts/benchmark - python benchmark.py + cd $BASE/MET/internal/scripts/benchmark + python benchmark.py - .. note:: + .. note:: The intermediate summary_output.txt and detail_output.txt files generated by CTRACK are found in the directory from which the benchmark.py script was invoked (in the $BASE/MET/internal/scripts/benchmark directory). @@ -477,7 +477,7 @@ Overview of Steps for Performing Benchmarking View the consolidated metrics to identify potential performance enhancements. Refer to the CTRACK documentation to learn about the metrics collected, under the **Metrics & Output** section: - https://github.com/Compaile/ctrack#metrics--output + https://github.com/Compaile/ctrack#metrics--output .. note:: @@ -517,11 +517,11 @@ Keywords .. note:: - - CTRACK - - benchmarking - - profiling - - code profiler - - code profiling + - CTRACK + - benchmarking + - profiling + - code profiler + - code profiling diff --git a/docs/Contributors_Guide/dev_details/index.rst b/docs/Contributors_Guide/dev_details/index.rst index fde3513711..318be0daee 100644 --- a/docs/Contributors_Guide/dev_details/index.rst +++ b/docs/Contributors_Guide/dev_details/index.rst @@ -6,7 +6,7 @@ This chapter provides specific details about select topics within the MET code base. The list of topics is certainly not comprehensive. .. toctree:: - :titlesonly: + :titlesonly: - tmp_file_use - static_data_files + tmp_file_use + static_data_files diff --git a/docs/Contributors_Guide/dev_details/static_data_files.rst b/docs/Contributors_Guide/dev_details/static_data_files.rst index 6bdb5cc48b..93928d4c52 100644 --- a/docs/Contributors_Guide/dev_details/static_data_files.rst +++ b/docs/Contributors_Guide/dev_details/static_data_files.rst @@ -45,7 +45,7 @@ recommended update frequency and method. :numref:`User's Guide Section %s `, is read by ASCII2NC and contains buoy latitude and longitude locations that can change on a daily basis. To be used in real time, this file should be - regenerated daily and the :code:`MET_NDBC_STATION` environment variable + regenerated daily and the :code:`MET_NDBC_STATIONS` environment variable should define its location. Use the :code:`scripts/python/utility/build_ndbc_stations_from_web.py` utility to update its contents. diff --git a/docs/Contributors_Guide/dev_details/tmp_file_use.rst b/docs/Contributors_Guide/dev_details/tmp_file_use.rst index 327066e655..77b7811d4b 100644 --- a/docs/Contributors_Guide/dev_details/tmp_file_use.rst +++ b/docs/Contributors_Guide/dev_details/tmp_file_use.rst @@ -5,7 +5,7 @@ Use of Temporary Files The MET application and library code uses temporary files in several places. Each specific use of temporary files is described below. The -directory in which temporary files are stored is configurable as, +directory in which temporary files are stored is configurable, as described in :numref:`User's Guide Section %s `. Whenever a MET application is run, the operating system assigns it a @@ -93,9 +93,9 @@ Where {LINE_TYPE} is :code:`cnt`, :code:`cts`, :code:`mcts`, :code:`nbrcnt`, or :code:`nbrcts`. .. note:: - Consider whether or not it's realistic to hold the resampled - statistics in memory rather than writing them to temporary files. - If so, that would reduce the I/O. + Consider whether or not it's realistic to hold the resampled + statistics in memory rather than writing them to temporary files. + If so, that would reduce the I/O. .. _tmp_files_stat_analysis: @@ -117,15 +117,15 @@ input data for each job. * :code:`tmp_stat_analysis_{PID}`: If warranted, Stat-Analysis reads all input data, applies common filtering logic, and writes the - result to this temporary file. All of analysis jobs read data from + result to this temporary file. All of the analysis jobs read data from this temporary file, apply any additional job-specific filtering criteria, and perform the requested operation. .. note:: - Earlier versions of Stat-Analysis always wrote a temporary file - regardless of the number of jobs and filtering criteria. That - logic has been refined to only use temporary files when they may - increase efficiency. + Earlier versions of Stat-Analysis always wrote a temporary file + regardless of the number of jobs and filtering criteria. That + logic has been refined to only use temporary files when they may + increase efficiency. .. _tmp_files_python_embedding: diff --git a/docs/Contributors_Guide/index.rst b/docs/Contributors_Guide/index.rst index c069b4e345..ad8140d5fb 100644 --- a/docs/Contributors_Guide/index.rst +++ b/docs/Contributors_Guide/index.rst @@ -5,21 +5,21 @@ Contributor's Guide Welcome to the Model Evaluation Tools (MET) Contributor's Guide. .. toctree:: - :titlesonly: - :numbered: - :maxdepth: 1 + :titlesonly: + :numbered: + :maxdepth: 1 - coding_standards - dev_env - dev_details/index - github_workflow - testing - continuous_integration - code_profiling - dockerhub - documentation - templates - user_support + coding_standards + dev_env + dev_details/index + github_workflow + testing + continuous_integration + code_profiling + dockerhub + documentation + templates + user_support Indices and tables ================== diff --git a/docs/Contributors_Guide/testing.rst b/docs/Contributors_Guide/testing.rst index 2849c29df5..4ebbf694c0 100644 --- a/docs/Contributors_Guide/testing.rst +++ b/docs/Contributors_Guide/testing.rst @@ -5,12 +5,12 @@ Testing make test ========= -After MET has been compiled, run ``make test`` from the top-level directory to execute the scripts found in the ``scripts/examples`` directory. These scripts run a subset of the MET tools reading input data from the top-level ``data`` directory and configuration files from the ``scripts/config`` directory and write output to top-level ``out`` directory. Successful completion of these tests provides reasonable assurance that MET has been compiled well and is running properly. However, these sample scripts are not comprehensive and do not exercise all possible configuration options. So it's possible for the ``make test`` scripts to run without error, but for users to later encounter issues when running MET with new inputs files and configuration options. +After MET has been compiled, run ``make test`` from the top-level directory to execute the scripts found in the ``scripts/examples`` directory. These scripts run a subset of the MET tools reading input data from the top-level ``data`` directory and configuration files from the ``scripts/config`` directory and write output to the top-level ``out`` directory. Successful completion of these tests provides reasonable assurance that MET has been compiled well and is running properly. However, these sample scripts are not comprehensive and do not exercise all possible configuration options. So it's possible for the ``make test`` scripts to run without error, but for users to later encounter issues when running MET with new input files and configuration options. Unit Tests ========== -The MET unit tests offer much more thorough testing coverage of the MET tools than running ``make test``, as described above. These units tests provide the basis for the regression testing performed for each pull request. Logic exists in GitHub automation to run these unit tests and check for differences in the output. However, these unit tests can also be run locally and instructions for doing so are provided in this section. +The MET unit tests offer much more thorough testing coverage of the MET tools than running ``make test``, as described above. These unit tests provide the basis for the regression testing performed for each pull request. Logic exists in GitHub automation to run these unit tests and check for differences in the output. However, these unit tests can also be run locally and instructions for doing so are provided in this section. Running Unit Tests ------------------ @@ -34,10 +34,10 @@ Set the required environment variables needed to run. Example:: - export MET_BASE=/path/to/install/MET/share/met - export MET_TEST_BASE=/path/to/src/MET/internal/test_unit - export MET_TEST_INPUT=/path/to/MET_unit_test - export MET_TEST_OUTPUT=/path/to/my/output_dir + export MET_BASE=/path/to/install/MET/share/met + export MET_TEST_BASE=/path/to/src/MET/internal/test_unit + export MET_TEST_INPUT=/path/to/MET_unit_test + export MET_TEST_OUTPUT=/path/to/my/output_dir Other environment variables required for some of the unit tests include: @@ -48,34 +48,34 @@ Run the tests Navigate to the *internal/test_unit* directory of the MET repository:: - cd ${MET_TEST_BASE} + cd ${MET_TEST_BASE} To run all of the unit tests, call the *bin/unit_test.sh* script:: - ./bin/unit_test.sh + ./bin/unit_test.sh To run a single unit test group, call the *python/unit.py* script, passing it an XML test config file:: - ./python/unit.py ./xml/unit_pcp_combine.xml + ./python/unit.py ./xml/unit_pcp_combine.xml To generate commands corresponding to a single unit test group, but not actually execute those commands, add the *-cmd* command line argument and redirect the output to a file:: - ./python/unit.py ./xml/unit_pcp_combine.xml -cmd > unit_pcp_combine.sh + ./python/unit.py ./xml/unit_pcp_combine.xml -cmd > unit_pcp_combine.sh Extracting individual commands to be executed in this way can be convenient during the software development process. .. note:: - Some unit tests depend on the output of other unit tests. - For example, *unit_plot_data_plane.xml* requires output from *unit_pcp_combine.xml*. - Those dependencies are generally noted in comments at the top of each unit test xml file. + Some unit tests depend on the output of other unit tests. + For example, *unit_plot_data_plane.xml* requires output from *unit_pcp_combine.xml*. + Those dependencies are generally noted in comments at the top of each unit test xml file. Input Data ---------- Input data used to run the MET unit tests in CI workflows are pulled from the DTC web server and stored on DockerHub. -On the web server, data is stored for each supported version, e.g. *v12.0*, *v12.1*, etc. +On the web server, data is stored for each supported version, e.g., *v12.0*, *v12.1*, etc. There is also a directory called *develop* that includes symbolic links to the latest version, which is the version that is currently in development. This is done so that the latest state of the input data is used for new development @@ -87,13 +87,13 @@ Setting up a new web server .. note:: - These instructions require access to run commands as the *met_test* user on the DTC web server. + These instructions require access to run commands as the *met_test* user on the DTC web server. The GitHub Actions custom action `metplus-action-data-update `_ expects a specific URL defined in *update_data_volumes.py* script in its repo. This directory should exist on the web server. -This can be a link to another directory, but the name must match the repo name, e.g. MET. +This can be a link to another directory, but the name must match the repo name, e.g., MET. If this path must differ on a new web server, then modifications will be needed to the custom action. The directory should also be linked from the *met_test* user's home directory with the name *MET_unit_test*. @@ -106,11 +106,11 @@ update the input data and set up the next release directory. This is not necessarily required, but makes it convenient to find and call the script. :: - runas met_test - cd ~/ - git clone --branch develop https://github.com/dtcenter/MET - ln -s MET/internal/scripts/unit_test_ci/setup_met_next_release_data.sh - ln -s MET/internal/scripts/unit_test_ci/update_met_unit_test_data.sh + runas met_test + cd ~/ + git clone --branch develop https://github.com/dtcenter/MET + ln -s MET/internal/scripts/unit_test_ci/setup_met_next_release_data.sh + ln -s MET/internal/scripts/unit_test_ci/update_met_unit_test_data.sh The unit test input data directory contains directories for *develop* and each *vX.Y* version that is supported. Each directory should contain a tarfile called **unit_test-all.tgz** and a file called **volume_mount_directories**. @@ -130,7 +130,7 @@ Setup next development cycle .. note:: - These instructions require access to run commands as the *met_test* user on the DTC web server. + These instructions require access to run commands as the *met_test* user on the DTC web server. Once the *main_vX.Y* branch for a release has been created, the *develop* branch will contain development towards the next release. At this time, a new set of test data should be created for the next @@ -138,18 +138,18 @@ release so that it can be updated while preserving the test data used for an off For example, if the *main_v12.1* branch was created when the *12.1.0-rc1* release was created, then a data directory to store data for *v13.0* (or similar) should be created. -Pull changes from develop to ensure that the latest version of script is used. +Pull changes from develop to ensure that the latest version of the script is used. :: - runas met_test - cd ~/MET - git checkout develop - git pull + runas met_test + cd ~/MET + git checkout develop + git pull Run the script, passing the *vX.Y* version of the next release as an argument. If the script is linked from the home directory, run:: - ~/setup_met_next_release_data.sh v13.0 + ~/setup_met_next_release_data.sh v13.0 This will create the *v13.0* directory, copy the latest tarfile and volume mount files into *v13.0*, extract the tarfile contents into the *v13.0*, and update the symbolic links in the *develop* directory @@ -160,49 +160,49 @@ Adding new test files .. note:: - These instructions require access to run commands as the *met_test* user on the DTC web server. + These instructions require access to run commands as the *met_test* user on the DTC web server. -Updates to the input data, e.g. adding new test files, are made on the DTC web server. +Updates to the input data, e.g., adding new test files, are made on the DTC web server. The next time the MET CI unit tests are run, the web server will be checked and the input data will be updated automatically. Note that the unit tests are only run for develop/main branches or running via workflow dispatch. A push event to a branch will not run the full unit test suite and therefore will not update the input data. In the *MET_unit_test* directory, there is a directory called *unit_test*. -These files are the full set of fields and fields used for the unit tests. +These files are the full set of input files used for the unit tests. **These files are used by the MET regression tests that are run locally.** First, add any new files to the *unit_test* directory so they will be available to the MET regression tests. Example:: - cp /path/to/my/file.ext MET_unit_test/unit_test/DIRNAME/ + cp /path/to/my/file.ext MET_unit_test/unit_test/DIRNAME/ Next, add the new input files in the *unit_test* directory under the *vX.Y* directory that corresponds to the current development cycle. Example:: - cp /path/to/my/file.ext MET_unit_test/v23.1/unit_test/DIRNAME/ + cp /path/to/my/file.ext MET_unit_test/v23.1/unit_test/DIRNAME/ If any of the files are very large, consider creating a subset of these files. For example, GRIB2 files can be subset with *wgrib2* and NetCDF files can be subset using NCO tools. After the updates have been made, run the script to update the test data tarfile. -Pull changes from develop to ensure that the latest version of script is used. +Pull changes from develop to ensure that the latest version of the script is used. :: - runas met_test - cd ~/MET - git checkout develop - git pull + runas met_test + cd ~/MET + git checkout develop + git pull Run the script, passing the *vX.Y* version of the next release as an argument. If the script is linked from the home directory, run:: - ~/update_met_unit_test_data.sh v13.0 + ~/update_met_unit_test_data.sh v13.0 -This will save a copy the input data tarfile with the current date in YYYYMMDD format in case it needs to be recovered, +This will save a copy of the input data tarfile with the current date in YYYYMMDD format in case it needs to be recovered, then create the tarfile using the contents of the *unit_test* directory. diff --git a/docs/Users_Guide/appendixA.rst b/docs/Users_Guide/appendixA.rst index 2eadbc3851..ea80b4fd40 100644 --- a/docs/Users_Guide/appendixA.rst +++ b/docs/Users_Guide/appendixA.rst @@ -12,256 +12,256 @@ File-IO Q. How do I improve the speed of MET tools using Gen-Vx-Mask? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - The main reason to run gen_vx_mask is to make the MET - statistics tools (e.g. point_stat, grid_stat, or ensemble_stat) run - faster. The verification masking regions in those tools can be specified - as Lat/Lon polyline files or the NetCDF output of gen_vx_mask. However, - determining which grid points are inside/outside a polyline region can be - slow if the polyline contains many points or the grid is dense. Running - gen_vx_mask once to create a binary mask is much more efficient than - recomputing the mask when each MET statistics tool is run. If the polyline - only contains a small number of points or the grid is sparse running - gen_vx_mask first would only save a second or two. + The main reason to run gen_vx_mask is to make the MET + statistics tools (e.g., point_stat, grid_stat, or ensemble_stat) run + faster. The verification masking regions in those tools can be specified + as Lat/Lon polyline files or the NetCDF output of gen_vx_mask. However, + determining which grid points are inside/outside a polyline region can be + slow if the polyline contains many points or the grid is dense. Running + gen_vx_mask once to create a binary mask is much more efficient than + recomputing the mask when each MET statistics tool is run. If the polyline + only contains a small number of points or the grid is sparse running + gen_vx_mask first would only save a second or two. Q. How do I use map_data? ^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - The MET repository includes several map data files. Users can modify which - map datasets are included in the plots created by modifying the - configuration files for those tools. The default map datasets are defined - by the map_data dictionary in the ConfigMapData file. + The MET repository includes several map data files. Users can modify which + map datasets are included in the plots created by modifying the + configuration files for those tools. The default map datasets are defined + by the map_data dictionary in the ConfigMapData file. - .. code-block:: none + .. code-block:: none - map_data = { + map_data = { - line_color = [ 25, 25, 25 ]; // rgb triple values, 0-255 - line_width = 0.5; - line_dash = ""; + line_color = [ 25, 25, 25 ]; // rgb triple values, 0-255 + line_width = 0.5; + line_dash = ""; - source = [ - { file_name = "MET_BASE/map/country_data"; }, - { file_name = "MET_BASE/map/usa_state_data"; }, - { file_name = "MET_BASE/map/major_lakes_data"; } - ]; - } + source = [ + { file_name = "MET_BASE/map/country_data"; }, + { file_name = "MET_BASE/map/usa_state_data"; }, + { file_name = "MET_BASE/map/major_lakes_data"; } + ]; + } - Users can modify the ConfigMapData contents prior to running - 'make install'. - This will change the default map data for all of the MET tools which plots. - Alternatively, users can copy/paste/modify the map_data dictionary into the - configuration file for a MET tool. For example, you could add map_data to - the end of the MODE configuration file to customize plots created by MODE. + Users can modify the ConfigMapData contents prior to running + 'make install'. + This will change the default map data for all of the MET tools which plot. + Alternatively, users can copy/paste/modify the map_data dictionary into the + configuration file for a MET tool. For example, you could add map_data to + the end of the MODE configuration file to customize plots created by MODE. - Here is an example of running plot_data_plane and specifying the map_data - in the configuration string on the command line: + Here is an example of running plot_data_plane and specifying the map_data + in the configuration string on the command line: - .. code-block:: none + .. code-block:: none - plot_data_plane - sample.grib china_tmp_2m_admin.ps \ - 'name="TMP"; level="Z2"; \ - map_data = { source = [ { file_name = \ - "${MET_BASE}/map/admin_by_country/admin_China_data"; } \ - ]; }' + plot_data_plane \ + sample.grib china_tmp_2m_admin.ps \ + 'name="TMP"; level="Z2"; \ + map_data = { source = [ { file_name = \ + "${MET_BASE}/map/admin_by_country/admin_China_data"; } \ + ]; }' Q. How can I understand the number of matched pairs? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer - - Statistics are computed on matched forecast/observation pairs data. - For example, if the dimension of the grid is 37x37 up to - 1369 matched pairs are possible. However, if the forecast or - observation contains bad data at a point, that matched pair would - not be included in the calculations. There are a number of reasons that - observations could be rejected - mismatches in station id, variable names, - valid times, bad values, data off the grid, etc. - For example, if the forecast field contains missing data around the - edge of the domain, then that is a reason there may be 992 matched pairs - instead of 1369. Users can use the ncview tool to look at an example - netCDF file or run their files through plot_data_plane to help identify - any potential issues. - - One common support question is "Why am I getting 0 matched pairs from - Point-Stat?". As mentioned above, there are many reasons why point - observations can be excluded from your analysis. If running point_stat with - at least verbosity level 2 (-v 2, the default value), zero matched pairs - will result in the following type of log messages to be printed: - - .. code-block:: none - - DEBUG 2: Processing TMP/Z2 versus TMP/Z2, for observation type ADPSFC, over region FULL, for interpolation method UW_MEAN(1), using 0 pairs. - DEBUG 2: Number of matched pairs = 0 - DEBUG 2: Observations processed = 1166 - DEBUG 2: Rejected: station id = 0 - DEBUG 2: Rejected: obs var name = 1166 - DEBUG 2: Rejected: valid time = 0 - DEBUG 2: Rejected: bad obs value = 0 - DEBUG 2: Rejected: off the grid = 0 - DEBUG 2: Rejected: topography = 0 - DEBUG 2: Rejected: level mismatch = 0 - DEBUG 2: Rejected: quality marker = 0 - DEBUG 2: Rejected: message type = 0 - DEBUG 2: Rejected: masking region = 0 - DEBUG 2: Rejected: bad fcst value = 0 - DEBUG 2: Rejected: bad climo mean = 0 - DEBUG 2: Rejected: bad climo stdev = 0 - DEBUG 2: Rejected: mpr filter = 0 - DEBUG 2: Rejected: duplicates = 0 - - This list of the rejection reason counts above matches the order in - which the filtering logic is applied in the code. In this example, - none of the point observations match the variable name requested - in the configuration file. So all of the 1166 observations are rejected - for the same reason. - - In addition, running point_stat with at least verbosity level 9 (-v 9) - will result in a log message being printed to explain why each - observation is skipped or retained for each verification task. - This level of detail is intended only for debugging purposes. +.. dropdown:: Answer + + Statistics are computed on matched forecast/observation pairs data. + For example, if the dimension of the grid is 37x37 up to + 1369 matched pairs are possible. However, if the forecast or + observation contains bad data at a point, that matched pair would + not be included in the calculations. There are a number of reasons that + observations could be rejected - mismatches in station id, variable names, + valid times, bad values, data off the grid, etc. + For example, if the forecast field contains missing data around the + edge of the domain, then that is a reason there may be 992 matched pairs + instead of 1369. Users can use the ncview tool to look at an example + NetCDF file or run their files through plot_data_plane to help identify + any potential issues. + + One common support question is "Why am I getting 0 matched pairs from + Point-Stat?". As mentioned above, there are many reasons why point + observations can be excluded from your analysis. If running point_stat with + at least verbosity level 2 (-v 2, the default value), zero matched pairs + will result in the following type of log messages to be printed: + + .. code-block:: none + + DEBUG 2: Processing TMP/Z2 versus TMP/Z2, for observation type ADPSFC, over region FULL, for interpolation method UW_MEAN(1), using 0 pairs. + DEBUG 2: Number of matched pairs = 0 + DEBUG 2: Observations processed = 1166 + DEBUG 2: Rejected: station id = 0 + DEBUG 2: Rejected: obs var name = 1166 + DEBUG 2: Rejected: valid time = 0 + DEBUG 2: Rejected: bad obs value = 0 + DEBUG 2: Rejected: off the grid = 0 + DEBUG 2: Rejected: topography = 0 + DEBUG 2: Rejected: level mismatch = 0 + DEBUG 2: Rejected: quality marker = 0 + DEBUG 2: Rejected: message type = 0 + DEBUG 2: Rejected: masking region = 0 + DEBUG 2: Rejected: bad fcst value = 0 + DEBUG 2: Rejected: bad climo mean = 0 + DEBUG 2: Rejected: bad climo stdev = 0 + DEBUG 2: Rejected: mpr filter = 0 + DEBUG 2: Rejected: duplicates = 0 + + This list of the rejection reason counts above matches the order in + which the filtering logic is applied in the code. In this example, + none of the point observations match the variable name requested + in the configuration file. So all of the 1166 observations are rejected + for the same reason. + + In addition, running point_stat with at least verbosity level 9 (-v 9) + will result in a log message being printed to explain why each + observation is skipped or retained for each verification task. + This level of detail is intended only for debugging purposes. Q. What types of NetCDF files can MET read? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - There are three flavors of NetCDF that MET can read directly. + There are three flavors of NetCDF that MET can read directly. - 1. Gridded NetCDF output from one of the MET tools + 1. Gridded NetCDF output from one of the MET tools - 2. Output from the WRF model that has been post-processed using - the wrf_interp utility + 2. Output from the WRF model that has been post-processed using + the wrf_interp utility - 3. NetCDF data following the `climate-forecast (CF) convention - `_ + 3. NetCDF data following the `climate-forecast (CF) convention + `_ - Lastly, users can write python scripts to pass data that's gridded to the - MET tools in memory. If the data doesn't fall into one of those categories, - then it's not a gridded dataset that MET can handle directly. - Satellite data, in general, will not be gridded. Typically it - contains a dense mesh of data at lat/lon points, but typically - those lat/lon points are not evenly spaced onto - a regular grid. + Lastly, users can write Python scripts to pass data that's gridded to the + MET tools in memory. If the data doesn't fall into one of those categories, + then it's not a gridded dataset that MET can handle directly. + Satellite data, in general, will not be gridded. Typically it + contains a dense mesh of data at lat/lon points, but typically + those lat/lon points are not evenly spaced onto + a regular grid. - While MET's point2grid tool does support some satellite data inputs, it is - limited. Using python embedding is another option for handling new datasets - not supported natively by MET. + While MET's point2grid tool does support some satellite data inputs, it is + limited. Using Python embedding is another option for handling new datasets + not supported natively by MET. Q. How do I choose a time slice in a NetCDF file? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - When processing NetCDF files, the level information needs to be - specified to tell MET which 2D slice of data to use. - The index is selected from - a value when it starts with "@" for vertical level (pressure or height) - and time. The actual time, @YYYYMMDD_HHMM, is allowed instead of selecting - the time index. + When processing NetCDF files, the level information needs to be + specified to tell MET which 2D slice of data to use. + The index is selected from + a value when it starts with "@" for vertical level (pressure or height) + and time. The actual time, @YYYYMMDD_HHMM, is allowed instead of selecting + the time index. - Let's use plot_data_plane as an example: + Let's use plot_data_plane as an example: - .. code-block:: none + .. code-block:: none - plot_data_plane \ - MERGE_20161201_20170228.nc \ - obs.ps \ - 'name="APCP"; level="(5,*,*)";' + plot_data_plane \ + MERGE_20161201_20170228.nc \ + obs.ps \ + 'name="APCP"; level="(5,*,*)";' - plot_data_plane \ - gtg_obs_forecast.20130730.i00.f00.nc \ - altitude_20000.ps \ - 'name = "edr"; level = "(@20130730_0000,@20000,*,*)";' + plot_data_plane \ + gtg_obs_forecast.20130730.i00.f00.nc \ + altitude_20000.ps \ + 'name = "edr"; level = "(@20130730_0000,@20000,*,*)";' - Assuming that the first array is the time, this will select the 6-th - time slice of the APCP data and plot it since these indices are 0-based. + Assuming that the first array is the time, this will select the 6-th + time slice of the APCP data and plot it since these indices are 0-based. Q. How do I use the UNIX time conversion? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - Regarding the timing information in the NetCDF variable attributes: + Regarding the timing information in the NetCDF variable attributes: - .. code-block:: none + .. code-block:: none - APCP_24:init_time_ut = 1306886400 ; + APCP_24:init_time_ut = 1306886400 ; - “ut” stands for UNIX time, which is the number of seconds - since Jan 1, 1970. It is a convenient way of storing timing - information since it is easy to add/subtract. The UNIX date command - can be used to convert back/forth between unix time and time strings: + “ut” stands for UNIX time, which is the number of seconds + since Jan 1, 1970. It is a convenient way of storing timing + information since it is easy to add/subtract. The UNIX date command + can be used to convert back/forth between unix time and time strings: - To convert unix time to ymd_hms date: + To convert unix time to ymd_hms date: - .. code-block:: none + .. code-block:: none - date -ud '1970-01-01 UTC '1306886400' seconds' +%Y%m%d_%H%M%S 20110601_000000 + date -ud '1970-01-01 UTC '1306886400' seconds' +%Y%m%d_%H%M%S 20110601_000000 - To convert ymd_hms to unix date: + To convert ymd_hms to unix date: - .. code-block:: none + .. code-block:: none - date -ud ''2011-06-01' UTC '00:00:00'' +%s 1306886400 + date -ud ''2011-06-01' UTC '00:00:00'' +%s 1306886400 - Regarding TRMM data, it may be easier to work with the binary data and - use the trmm2nc.R script described on this - `page `_ - under observation datasets. + Regarding TRMM data, it may be easier to work with the binary data and + use the trmm2nc.R script described on this + `page `_ + under observation datasets. - Follow the TRMM binary links to either the 3 or 24-hour accumulations, - save the files, and run them through that script. That is faster - and easier than trying to get an ASCII dump. That Rscript can also - subset the TRMM data if needed. Look for the section of it titled - "Output domain specification" and define the lat/lon's that needs - to be included in the output. + Follow the TRMM binary links to either the 3 or 24-hour accumulations, + save the files, and run them through that script. That is faster + and easier than trying to get an ASCII dump. That Rscript can also + subset the TRMM data if needed. Look for the section of it titled + "Output domain specification" and define the lat/lon's that need + to be included in the output. -Q. Does MET use a fixed-width output format for its ASCII output files? +Q. Does MET use a fixed-width output format for its ASCII output files? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - MET does not use the Fortran-like fixed width format in its - ASCII output file. Instead, the column widths are adjusted for each - run to insert at least one space between adjacent columns. The header - columns of the MET output contain user-defined strings which may be - of arbitrary length. For example, columns such as MODEL, OBTYPE, and - DESC may be set by the user to any string value. Additionally, the - amount of precision written is also configurable. The - "output_precision" config file entry can be changed from its default - value of 5 decimal places to up to 12 decimal places, which would also - impact the column widths of the output. + MET does not use the Fortran-like fixed width format in its + ASCII output file. Instead, the column widths are adjusted for each + run to insert at least one space between adjacent columns. The header + columns of the MET output contain user-defined strings which may be + of arbitrary length. For example, columns such as MODEL, OBTYPE, and + DESC may be set by the user to any string value. Additionally, the + amount of precision written is also configurable. The + "output_precision" config file entry can be changed from its default + value of 5 decimal places to up to 12 decimal places, which would also + impact the column widths of the output. - Due to these issues, it is not possible to select a reasonable fixed - width for each column ahead of time. The AsciiTable class in MET does - a lot of work to line up the output columns, to make sure there is - at least one space between them. + Due to these issues, it is not possible to select a reasonable fixed + width for each column ahead of time. The AsciiTable class in MET does + a lot of work to line up the output columns, to make sure there is + at least one space between them. - If a fixed-width format is needed, the easiest option would be - writing a script to post-process the MET output into the fixed-width - format that is needed or that the code expects. + If a fixed-width format is needed, the easiest option would be + writing a script to post-process the MET output into the fixed-width + format that is needed or that the code expects. Q. Do the ASCII output files created by MET use scientific notation? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - By default, the ASCII output files created by MET make use of - scientific notation when appropriate. The formatting of the - numbers that the AsciiTable class writes is handled by a call - to printf. The "%g" formatting option can result in - scientific notation: - http://www.cplusplus.com/reference/cstdio/printf/ + By default, the ASCII output files created by MET make use of + scientific notation when appropriate. The formatting of the + numbers that the AsciiTable class writes is handled by a call + to printf. The "%g" formatting option can result in + scientific notation: + http://www.cplusplus.com/reference/cstdio/printf/ - It has been recommended that a configuration option be added to - MET to disable the use of scientific notation. That enhancement - is planned for a future release. + It has been recommended that a configuration option be added to + MET to disable the use of scientific notation. That enhancement + is planned for a future release. Gen-Vx-Mask ----------- @@ -269,69 +269,69 @@ Gen-Vx-Mask Q. I have a list of stations to use for verification. I also have a poly region defined. If I specify both of these should the result be a union of them? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - These settings are defined in the "mask" section of the Point-Stat - configuration file. You can define masking regions in one of 3 ways, - as a "grid", a "poly" line file, or a "sid" list of station ID's. + These settings are defined in the "mask" section of the Point-Stat + configuration file. You can define masking regions in one of 3 ways, + as a "grid", a "poly" line file, or a "sid" list of station IDs. - If you specify one entry for "poly" and one entry for "sid", you - should see output for those two different masks. Note that each of - these settings is an array of values, as indicated by the square - brackets "[]" in the default config file. If you specify 5 grids, - 3 poly's, and 2 SID lists, you'd get output for those 10 separate - masking regions. Point-Stat does not compute unions or intersections - of masking regions. Instead, they are each processed separately. + If you specify one entry for "poly" and one entry for "sid", you + should see output for those two different masks. Note that each of + these settings is an array of values, as indicated by the square + brackets "[]" in the default config file. If you specify 5 grids, + 3 poly's, and 2 SID lists, you'd get output for those 10 separate + masking regions. Point-Stat does not compute unions or intersections + of masking regions. Instead, they are each processed separately. - Is it true that you really want to use a polyline to define an area - and then use a SID list to capture additional points outside of - that polyline? + Is it true that you really want to use a polyline to define an area + and then use a SID list to capture additional points outside of + that polyline? - If so, your options are: + If so, your options are: - 1. Define one single SID list which include all the points currently - inside the polyline as well as the extra ones outside. + 1. Define one single SID list which includes all the points currently + inside the polyline as well as the extra ones outside. - 2. Continue verifying using one polyline and one SID list and - write partial sums and contingency table counts. + 2. Continue verifying using one polyline and one SID list and + write partial sums and contingency table counts. - Then aggregate the results together by running a Stat-Analysis job. + Then aggregate the results together by running a Stat-Analysis job. Q. How do I define a masking region with a GFS file? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - Grab a sample GFS file: + Grab a sample GFS file: - .. code-block:: none + .. code-block:: none - wget - http://www.ftp.ncep.noaa.gov/data/nccf/com/gfs/prod/gfs/2016102512/gfs.t12z.pgrb2.0p50.f000 + wget + http://www.ftp.ncep.noaa.gov/data/nccf/com/gfs/prod/gfs/2016102512/gfs.t12z.pgrb2.0p50.f000 - Use the MET regrid_data_plane tool to put some data on a - lat/lon grid over Europe: + Use the MET regrid_data_plane tool to put some data on a + lat/lon grid over Europe: - .. code-block:: none + .. code-block:: none - regrid_data_plane gfs.t12z.pgrb2.0p50.f000 \ - 'latlon 100 100 25 0 0.5 0.5' gfs_euro.nc -field 'name="TMP"; level="Z2";' + regrid_data_plane gfs.t12z.pgrb2.0p50.f000 \ + 'latlon 100 100 25 0 0.5 0.5' gfs_euro.nc -field 'name="TMP"; level="Z2";' - Run the MET gen_vx_mask tool to apply your polyline to the European domain: + Run the MET gen_vx_mask tool to apply your polyline to the European domain: - .. code-block:: none + .. code-block:: none - gen_vx_mask gfs_euro.nc POLAND.poly POLAND_mask.nc + gen_vx_mask gfs_euro.nc POLAND.poly POLAND_mask.nc -type poly - Run the MET plot_data_plane tool to display the resulting mask field: + Run the MET plot_data_plane tool to display the resulting mask field: - .. code-block:: none + .. code-block:: none - plot_data_plane POLAND_mask.nc POLAND_mask.ps 'name="POLAND"; level="(*,*)";' + plot_data_plane POLAND_mask.nc POLAND_mask.ps 'name="POLAND"; level="(*,*)";' - In this example, the mask is in roughly the right spot, but there - are obvious problems with the latitude and longitude values used - to define that mask for Poland. + In this example, the mask is in roughly the right spot, but there + are obvious problems with the latitude and longitude values used + to define that mask for Poland. Grid-Stat --------- @@ -339,286 +339,287 @@ Grid-Stat Q. How do I define a complex masking region? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - A user can define intersections and unions of multiple fields to - define masks. - Prior to running Grid-Stat, the user can run the Gen-VX-Mask tool one or - more times to define a more complex masking area by thresholding multiple - fields. + A user can define intersections and unions of multiple fields to + define masks. + Prior to running Grid-Stat, the user can run the Gen-Vx-Mask tool one or + more times to define a more complex masking area by thresholding multiple + fields. - For example, using a forecast GRIB file (fcst.grb) which contains - 2 records, one for 2-m temperature and a second for 6-hr - accumulated precip. The only - grid points that are desired are grid points below freezing with non-zero - precip. The user should run Gen-Vx-Mask twice - once to define the - temperature mask and a second time to intersect that with the precip mask: + For example, using a forecast GRIB file (fcst.grb) which contains + 2 records, one for 2-m temperature and a second for 6-hr + accumulated precip. The only + grid points that are desired are grid points below freezing with non-zero + precip. The user should run Gen-Vx-Mask twice - once to define the + temperature mask and a second time to intersect that with the precip mask: - .. code-block:: none + .. code-block:: none + + gen_vx_mask fcst.grb fcst.grb tmp_mask.nc \ + -type data \ + -mask_field 'name="TMP"; level="Z2"' -thresh le273 - gen_vx_mask fcst.grb fcst.grb tmp_mask.nc \ - -type data \ - -mask_field 'name="TMP"; level="Z2"' -thresh le273 - gen_vx_mask tmp_mask.nc fcst.grb tmp_and_precip_mask.nc \ - -type data \ - -input_field 'name="TMP_Z2"; level="(*,*)";' \ - -mask_field 'name="APCP"; level="A6";' -thresh gt0 \ - -intersection -name "FREEZING_PRECIP" + gen_vx_mask tmp_mask.nc fcst.grb tmp_and_precip_mask.nc \ + -type data \ + -input_field 'name="TMP_Z2"; level="(*,*)";' \ + -mask_field 'name="APCP"; level="A6";' -thresh gt0 \ + -intersection -name "FREEZING_PRECIP" - The first one is pretty straight-forward. + The first one is pretty straightforward. - 1. The input field (fcst.grb) defines the domain for the mask. + 1. The input field (fcst.grb) defines the domain for the mask. - 2. Since we're doing data masking and the data we want lives in - fcst.grb, we pass it in again as the mask_file. + 2. Since we're doing data masking and the data we want lives in + fcst.grb, we pass it in again as the mask_file. - 3. Lastly "-mask_field" specifies the data we want from the mask file - and "-thresh" specifies the event threshold. + 3. Lastly "-mask_field" specifies the data we want from the mask file + and "-thresh" specifies the event threshold. - The second call is a bit tricky. + The second call is a bit tricky. - 1. Do data masking (-type data) + 1. Do data masking (-type data) - 2. Read the NetCDF variable named "TMP_Z2" from the input file - (tmp_mask.nc) + 2. Read the NetCDF variable named "TMP_Z2" from the input file + (tmp_mask.nc) - 3. Define the mask by reading 6-hour precip from the mask file - (fcst.grb) and looking for values > 0 (-mask_field) + 3. Define the mask by reading 6-hour precip from the mask file + (fcst.grb) and looking for values > 0 (-mask_field) - 4. Apply intersection logic when combining the "input" value with - the "mask" value (-intersection). + 4. Apply intersection logic when combining the "input" value with + the "mask" value (-intersection). - 5. Name the output NetCDF variable as "FREEZING_PRECIP" (-name). - This is totally optional, but convenient. + 5. Name the output NetCDF variable as "FREEZING_PRECIP" (-name). + This is totally optional, but convenient. - A user can write a script with multiple calls to Gen-Vx-Mask to - apply complex masking logic and then pass the output mask file - to Grid-Stat in its configuration file. + A user can write a script with multiple calls to Gen-Vx-Mask to + apply complex masking logic and then pass the output mask file + to Grid-Stat in its configuration file. Q. How do I use neighborhood methods to compute fraction skill score? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer - - A common application of fraction skill score (FSS) is comparing forecast - and observed thunderstorms. When computing FSS, first threshold the fields - to define events and non-events. Then look at successively larger and - larger areas around each grid point to see how the forecast event frequency - compares to the observed event frequency. - - Applying this method to rainfall (and monsoons) is also reasonable. - Keep in mind that Grid-Stat is the tool that computes FSS. Grid-Stat will - need to be run once for each evaluation time. As an example, evaluating - once per day, run Grid-Stat 122 times for the 122 days of a monsoon season. - This will result in 122 FSS values. These can be viewed as a time series, - or the Stat-Analysis tool could be used to aggregate them together into - a single FSS value, like this: - - .. code-block:: none - - stat_analysis -job aggregate -line_type NBRCNT \ - -lookin out/grid_stat - - Be sure to pick thresholds (e.g. for the thunderstorms and monsoons) - that capture the "events" that are of interest in studying. - - Also be aware that MET uses the "vld_thresh" setting in the configuration - file to decide how to handle data along the edge of the domain. Let us say - it is computing a fractional coverage field using a 5x5 neighborhood - and it is at the edge of the domain. 15 points contain valid data and - 10 points are outside the domain. Grid-Stat computes the valid data ratio - as 15/25 = 0.6. Then it applies the valid data threshold. Suppose - vld_thresh = 0.5. Since 0.6 > 0.5 MET will compute a fractional coverage - value for that point using the 15 valid data points. Next suppose - vld_thresh = 1.0. Since 0.6 is less than 1.0, MET will just skip that - point by setting it to bad data. - - Setting vld_thresh = 1.0 will ensure that FSS will only be computed at - points where all NxN values contain valid data. Setting it to 0.5 only - requires half of them. - -Q. Is an example of verifying forecast probabilities? -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - - .. dropdown:: Answer - - There is an example of verifying probabilities in the test scripts - included with the MET release. Take a look in: - - .. code-block:: none - - MET/scripts/config/GridStatConfig_POP_12 - - The config file should look something like this: - - .. code-block:: none - - fcst = { - wind_thresh = [ NA ]; - field = [ - { - name = "LCDC"; - level = [ "L0" ]; - prob = TRUE; - cat_thresh = [ >=0.0, >=0.1, >=0.2, >=0.3, >=0.4, >=0.5, >=0.6, >=0.7, >=0.8, >=0.9]; - } - ]; - }; - - obs = { - wind_thresh = [ NA ]; - field = [ - { - name = "WIND"; - level = [ "Z2" ]; - cat_thresh = [ >=34 ]; - } - ]; - }; - - The PROB flag is set to TRUE to tell grid_stat to process this as - probability data. The cat_thresh is set to partition the probability - values between 0 and 1. Note that if the probability data contains - values from 0 to 100, MET automatically divides by 100 to rescale to - the 0 to 1 range. +.. dropdown:: Answer + + A common application of fraction skill score (FSS) is comparing forecast + and observed thunderstorms. When computing FSS, first threshold the fields + to define events and non-events. Then look at successively larger and + larger areas around each grid point to see how the forecast event frequency + compares to the observed event frequency. + + Applying this method to rainfall (and monsoons) is also reasonable. + Keep in mind that Grid-Stat is the tool that computes FSS. Grid-Stat will + need to be run once for each evaluation time. As an example, evaluating + once per day, run Grid-Stat 122 times for the 122 days of a monsoon season. + This will result in 122 FSS values. These can be viewed as a time series, + or the Stat-Analysis tool could be used to aggregate them together into + a single FSS value, like this: + + .. code-block:: none + + stat_analysis -job aggregate -line_type NBRCNT \ + -lookin out/grid_stat + + Be sure to pick thresholds (e.g., for the thunderstorms and monsoons) + that capture the "events" that are of interest in studying. + + Also be aware that MET uses the "vld_thresh" setting in the configuration + file to decide how to handle data along the edge of the domain. Let us say + it is computing a fractional coverage field using a 5x5 neighborhood + and it is at the edge of the domain. 15 points contain valid data and + 10 points are outside the domain. Grid-Stat computes the valid data ratio + as 15/25 = 0.6. Then it applies the valid data threshold. Suppose + vld_thresh = 0.5. Since 0.6 > 0.5 MET will compute a fractional coverage + value for that point using the 15 valid data points. Next suppose + vld_thresh = 1.0. Since 0.6 is less than 1.0, MET will just skip that + point by setting it to bad data. + + Setting vld_thresh = 1.0 will ensure that FSS will only be computed at + points where all NxN values contain valid data. Setting it to 0.5 only + requires half of them. + +Q. Is there an example of verifying forecast probabilities? +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +.. dropdown:: Answer + + There is an example of verifying probabilities in the test scripts + included with the MET release. Take a look in: + + .. code-block:: none + + MET/scripts/config/GridStatConfig_POP_12 + + The config file should look something like this: + + .. code-block:: none + + fcst = { + wind_thresh = [ NA ]; + field = [ + { + name = "LCDC"; + level = [ "L0" ]; + prob = TRUE; + cat_thresh = [ >=0.0, >=0.1, >=0.2, >=0.3, >=0.4, >=0.5, >=0.6, >=0.7, >=0.8, >=0.9]; + } + ]; + }; + + obs = { + wind_thresh = [ NA ]; + field = [ + { + name = "WIND"; + level = [ "Z2" ]; + cat_thresh = [ >=34 ]; + } + ]; + }; + + The PROB flag is set to TRUE to tell grid_stat to process this as + probability data. The cat_thresh is set to partition the probability + values between 0 and 1. Note that if the probability data contains + values from 0 to 100, MET automatically divides by 100 to rescale to + the 0 to 1 range. Q. What is an example of using Grid-Stat with regridding and masking turned on? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - Run Grid-Stat using the following commands and the attached config file + Run Grid-Stat using the following commands and the attached config file - .. code-block:: none + .. code-block:: none - mkdir out - grid_stat \ - gfs_4_20160220_0000_012.grb2 \ - ST4.2016022012.06h \ - GridStatConfig \ - -outdir out + mkdir out + grid_stat \ + gfs_4_20160220_0000_012.grb2 \ + ST4.2016022012.06h \ + GridStatConfig \ + -outdir out - Note the following two sections of the Grid-Stat config file: + Note the following two sections of the Grid-Stat config file: - .. code-block:: none + .. code-block:: none - regrid = { - to_grid = OBS; - vld_thresh = 0.5; - method = BUDGET; - width = 2; - } - - This tells Grid-Stat to do verification on the "observation" grid. - Grid-Stat reads the GFS and Stage4 data and then automatically regrids - the GFS data to the Stage4 domain using budget interpolation. - Use FCST to verify the forecast domain. And use either a named - grid or a grid specification string to regrid both the forecast and - observation to a common grid. For example, to_grid = "G212"; will - regrid both to NCEP Grid 212 before comparing them. + regrid = { + to_grid = OBS; + vld_thresh = 0.5; + method = BUDGET; + width = 2; + } - .. code-block:: none + This tells Grid-Stat to do verification on the "observation" grid. + Grid-Stat reads the GFS and Stage4 data and then automatically regrids + the GFS data to the Stage4 domain using budget interpolation. + Use FCST to verify the forecast domain. And use either a named + grid or a grid specification string to regrid both the forecast and + observation to a common grid. For example, to_grid = "G212"; will + regrid both to NCEP Grid 212 before comparing them. - mask = { grid = [ "FULL" ]; - poly = [ "MET_BASE/poly/CONUS.poly" ]; } + .. code-block:: none - This will compute statistics over the FULL model domain as well - as the CONUS masking area. + mask = { grid = [ "FULL" ]; + poly = [ "MET_BASE/poly/CONUS.poly" ]; } - To demonstrate that Grid-Stat worked as expected, run the following - commands to plot its NetCDF matched pairs output file: + This will compute statistics over the FULL model domain as well + as the CONUS masking area. - .. code-block:: none + To demonstrate that Grid-Stat worked as expected, run the following + commands to plot its NetCDF matched pairs output file: + + .. code-block:: none - plot_data_plane \ - out/grid_stat_120000L_20160220_120000V_pairs.nc \ - out/DIFF_APCP_06_A06_APCP_06_A06_CONUS.ps \ - 'name="DIFF_APCP_06_A06_APCP_06_A06_CONUS"; level="(*,*)";' + plot_data_plane \ + out/grid_stat_120000L_20160220_120000V_pairs.nc \ + out/DIFF_APCP_06_A06_APCP_06_A06_CONUS.ps \ + 'name="DIFF_APCP_06_A06_APCP_06_A06_CONUS"; level="(*,*)";' - Examine the resulting plot of that difference field. + Examine the resulting plot of that difference field. - Lastly, there is another option for defining that masking region. - Rather than passing the ascii CONUS.poly file to grid_stat, run the - gen_vx_mask tool and pass the NetCDF output of that tool to grid_stat. - The advantage to gen_vx_mask is that it will make grid_stat run a - bit faster. It can be used to construct much more complex masking areas. + Lastly, there is another option for defining that masking region. + Rather than passing the ASCII CONUS.poly file to grid_stat, run the + gen_vx_mask tool and pass the NetCDF output of that tool to grid_stat. + The advantage to gen_vx_mask is that it will make grid_stat run a + bit faster. It can be used to construct much more complex masking areas. Q. How do I use one mask for the forecast field and a different mask for the observation field? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - You can't define different - masks for the forecast and observation fields in MET tools. - MET only lets you - define a single mask (a masking grid or polyline) and then you choose - whether you want to apply it to the FCST, OBS, or BOTH of them. + You can't define different + masks for the forecast and observation fields in MET tools. + MET only lets you + define a single mask (a masking grid or polyline) and then you choose + whether you want to apply it to the FCST, OBS, or BOTH of them. - Nonetheless, there is a way you can accomplish this logic using the - gen_vx_mask tool. You run it once to pre-process the forecast field - and a second time to pre-process the observation field. And then pass - those output files to your desired MET tool. + Nonetheless, there is a way you can accomplish this logic using the + gen_vx_mask tool. You run it once to pre-process the forecast field + and a second time to pre-process the observation field. And then pass + those output files to your desired MET tool. - Below is an example using sample data that is included with the MET - release tarball. To illustrate, this command will read 3-hour - precip and 2-meter temperature, and resets the precip at any grid - point where the temperature is less than 290 K to a value of 0: + Below is an example using sample data that is included with the MET + release tarball. To illustrate, this command will read 3-hour + precip and 2-meter temperature, and reset the precip at any grid + point where the temperature is less than 290 K to a value of 0: - .. code-block:: none + .. code-block:: none - gen_vx_mask \ - data/sample_fcst/2005080700/wrfprs_ruc13_12.tm00_G212 \ - data/sample_fcst/2005080700/wrfprs_ruc13_12.tm00_G212 \ - APCP_03_where_2m_TMPge290.nc \ - -type data \ - -input_field 'name="APCP"; level="A3";' \ - -mask_field 'name="TMP"; level="Z2";' \ - -thresh 'lt290&&ne-9999' -v 4 -value 0 + gen_vx_mask \ + data/sample_fcst/2005080700/wrfprs_ruc13_12.tm00_G212 \ + data/sample_fcst/2005080700/wrfprs_ruc13_12.tm00_G212 \ + APCP_03_where_2m_TMPge290.nc \ + -type data \ + -input_field 'name="APCP"; level="A3";' \ + -mask_field 'name="TMP"; level="Z2";' \ + -thresh 'lt290&&ne-9999' -v 4 -value 0 - So this is a bit confusing. Here's what is happening: + So this is a bit confusing. Here's what is happening: - * The first argument is the input file which defines the grid. + * The first argument is the input file which defines the grid. - * The second argument is used to define the masking region and - since I'm reading data from the same input file, I've listed - that file twice. + * The second argument is used to define the masking region and + since I'm reading data from the same input file, I've listed + that file twice. - * The third argument is the output file name. + * The third argument is the output file name. - * The type of masking is "data" masking where we read a 2D field of - data and apply a threshold. + * The type of masking is "data" masking where we read a 2D field of + data and apply a threshold. - * By default, gen_vx_mask initializes each grid point to a value - of 0. Specifying "-input_field" tells it to initialize each grid - point to the value of that field (in my example 3-hour precip). + * By default, gen_vx_mask initializes each grid point to a value + of 0. Specifying "-input_field" tells it to initialize each grid + point to the value of that field (in my example 3-hour precip). - * The "-mask_field" option defines the data field that should be - thresholded. + * The "-mask_field" option defines the data field that should be + thresholded. - * The "-thresh" option defines the threshold to be applied. + * The "-thresh" option defines the threshold to be applied. - * The "-value" option tells it what "mask" value to write to the - output, and I've chosen 0. + * The "-value" option tells it what "mask" value to write to the + output, and I've chosen 0. - The example threshold is less than 290 and not -9999 (which is MET's - internal missing data value). So any grid point where the 2 meter - temperature is less than 290 K and is not bad data will be replaced - by a value of 0. + The example threshold is less than 290 and not -9999 (which is MET's + internal missing data value). So any grid point where the 2 meter + temperature is less than 290 K and is not bad data will be replaced + by a value of 0. - To more easily demonstrate this, I changed to using "-value 10" and ran - the output through plot_data_plane: + To more easily demonstrate this, I changed to using "-value 10" and ran + the output through plot_data_plane: - .. code-block:: none + .. code-block:: none - plot_data_plane \ - APCP_03_where_2m_TMPge290.nc \ - APCP_03_where_2m_TMPge290.ps \ - 'name="data_mask"; level="(*,*)";' + plot_data_plane \ + APCP_03_where_2m_TMPge290.nc \ + APCP_03_where_2m_TMPge290.ps \ + 'name="data_mask"; level="(*,*)";' - In the resulting plot, anywhere you see the pink value of 10, that's - where gen_vx_mask has masked out the grid point. + In the resulting plot, anywhere you see the pink value of 10, that's + where gen_vx_mask has masked out the grid point. Pcp-Combine ----------- @@ -626,384 +627,384 @@ Pcp-Combine Q. How do I add and subtract with Pcp-Combine? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - An example of running the MET pcp_combine tool to put NAM 3-hourly - precipitation accumulations data into user-desired 3 hour intervals is - provided below. + An example of running the MET pcp_combine tool to put NAM 3-hourly + precipitation accumulations data into user-desired 3 hour intervals is + provided below. - If the user wanted a 0-3 hour accumulation, this is already available - in the 03 UTC file. Run this file - through pcp_combine as a pass-through to put it into NetCDF format: + If the user wanted a 0-3 hour accumulation, this is already available + in the 03 UTC file. Run this file + through pcp_combine as a pass-through to put it into NetCDF format: - .. code-block:: none + .. code-block:: none - pcp_combine -add 03_file.grb 03 APCP_00_03.nc + pcp_combine -add 03_file.grb 03 APCP_00_03.nc - If the user wanted the 3-6 hour accumulation, they would subtract - 0-6 and 0-3 accumulations: + If the user wanted the 3-6 hour accumulation, they would subtract + 0-6 and 0-3 accumulations: - .. code-block:: none + .. code-block:: none - pcp_combine -subtract 06_file.grb 06 03_file.grb 03 APCP_03_06.nc + pcp_combine -subtract 06_file.grb 06 03_file.grb 03 APCP_03_06.nc - Similarly, if they wanted the 6-9 hour accumulation, they would - subtract 0-9 and 0-6 accumulations: + Similarly, if they wanted the 6-9 hour accumulation, they would + subtract 0-9 and 0-6 accumulations: - .. code-block:: none + .. code-block:: none - pcp_combine -subtract 09_file.grb 09 06_file.grb 06 APCP_06_09.nc + pcp_combine -subtract 09_file.grb 09 06_file.grb 06 APCP_06_09.nc - And so on. + And so on. - Run the 0-3 and 12-15 through pcp_combine even though they already have - the 3-hour accumulation. That way, all of the NAM files will be in the - same file format, and can use the same configuration file settings for - the other MET tools (grid_stat, mode, etc.). If the NAM files are a mix - of GRIB and NetCDF, the logic would need to be a bit more complicated. + Run the 0-3 and 12-15 through pcp_combine even though they already have + the 3-hour accumulation. That way, all of the NAM files will be in the + same file format, and can use the same configuration file settings for + the other MET tools (grid_stat, mode, etc.). If the NAM files are a mix + of GRIB and NetCDF, the logic would need to be a bit more complicated. Q. How do I combine 12-hour accumulated precipitation from two different initialization times? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - The "-sum" command assumes the same initialization time. Use the "-add" - option instead. + The "-sum" command assumes the same initialization time. Use the "-add" + option instead. - .. code-block:: none + .. code-block:: none - pcp_combine -add \ - WRFPRS_1997-06-03_APCP_A12.nc 'name="APCP_12"; level="(*,*)";' \ - WRFPRS_d01_1997-06-04_00_APCP_A12.grb 12 \ - Sum.nc + pcp_combine -add \ + WRFPRS_1997-06-03_APCP_A12.nc 'name="APCP_12"; level="(*,*)";' \ + WRFPRS_d01_1997-06-04_00_APCP_A12.grb 12 \ + Sum.nc - For the first file, list the file name followed by a config string - describing the field to use from the NetCDF file. For the second file, - list the file name followed by the accumulation interval to use - (12 for 12 hours). The output file, Sum.nc, will contain the - combine 12-hour accumulated precipitation. + For the first file, list the file name followed by a config string + describing the field to use from the NetCDF file. For the second file, + list the file name followed by the accumulation interval to use + (12 for 12 hours). The output file, Sum.nc, will contain the + combined 24-hour accumulated precipitation. - Here is a small excerpt from the pcp_combine usage statement: + Here is a small excerpt from the pcp_combine usage statement: - Note: For “-add” and "-subtract”, the accumulation intervals may be - substituted with config file strings. For that first file, we replaced - the accumulation interval with a config file string. + Note: For "-add" and "-subtract", the accumulation intervals may be + substituted with config file strings. For that first file, we replaced + the accumulation interval with a config file string. - Here are 3 commands you could use to plot these data files: + Here are 3 commands you could use to plot these data files: - .. code-block:: none + .. code-block:: none - plot_data_plane WRFPRS_1997-06-03_APCP_A12.nc \ - WRFPRS_1997-06-03_APCP_A12.ps 'name="APCP_12"; level="(*,*)";' + plot_data_plane WRFPRS_1997-06-03_APCP_A12.nc \ + WRFPRS_1997-06-03_APCP_A12.ps 'name="APCP_12"; level="(*,*)";' - .. code-block:: none + .. code-block:: none - plot_data_plane WRFPRS_d01_1997-06-04_00_APCP_A12.grb \ - WRFPRS_d01_1997-06-04_00_APCP_A12.ps 'name="APCP" level="A12";' + plot_data_plane WRFPRS_d01_1997-06-04_00_APCP_A12.grb \ + WRFPRS_d01_1997-06-04_00_APCP_A12.ps 'name="APCP"; level="A12";' - .. code-block:: none + .. code-block:: none - plot_data_plane sum.nc sum.ps 'name="APCP_24"; level="(*,*)";' + plot_data_plane Sum.nc Sum.ps 'name="APCP_24"; level="(*,*)";' Q. How do I correct a precipitation time range? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - Typically, accumulated precipitation is stored in GRIB files using an - accumulation interval with a "time range" indicator value of 4. Here is - a description of the different time range indicator values and - meanings: http://www.nco.ncep.noaa.gov/pmb/docs/on388/table5.html + Typically, accumulated precipitation is stored in GRIB files using an + accumulation interval with a "time range" indicator value of 4. Here is + a description of the different time range indicator values and + meanings: http://www.nco.ncep.noaa.gov/pmb/docs/on388/table5.html - For example, take a look at the APCP in the GRIB files included in the - MET tar ball: + For example, take a look at the APCP in the GRIB files included in the + MET tar ball: - .. code-block:: none + .. code-block:: none - wgrib MET/data/sample_fcst/2005080700/wrfprs_ruc13_12.tm00_G212 | grep APCP - 1:0:d=05080700:APCP:kpds5=61:kpds6=1:kpds7=0:TR=4:P1=0: \ - P2=12:TimeU=1:sfc:0- 12hr acc:NAve=0 - 2:31408:d=05080700:APCP:kpds5=61:kpds6=1:kpds7=0:TR=4: \ - P1=9:P2=12:TimeU=1:sfc:9- 12hr acc:NAve=0 + wgrib MET/data/sample_fcst/2005080700/wrfprs_ruc13_12.tm00_G212 | grep APCP + 1:0:d=05080700:APCP:kpds5=61:kpds6=1:kpds7=0:TR=4:P1=0: \ + P2=12:TimeU=1:sfc:0- 12hr acc:NAve=0 + 2:31408:d=05080700:APCP:kpds5=61:kpds6=1:kpds7=0:TR=4: \ + P1=9:P2=12:TimeU=1:sfc:9- 12hr acc:NAve=0 - The "TR=4" indicates that these records contain an accumulation - between times P1 and P2. In the first record, the precip is accumulated - between 0 and 12 hours. In the second record, the precip is accumulated - between 9 and 12 hours. + The "TR=4" indicates that these records contain an accumulation + between times P1 and P2. In the first record, the precip is accumulated + between 0 and 12 hours. In the second record, the precip is accumulated + between 9 and 12 hours. - However, the GRIB data uses a time range indicator of 5, not 4. + However, the GRIB data uses a time range indicator of 5, not 4. - .. code-block:: none + .. code-block:: none - wgrib rmf_gra_2016040600.24 | grep APCP - 291:28360360:d=16040600:APCP:kpds5=61:kpds6=1:kpds7=0: \ - TR=5:P1=0:P2=24:TimeU=1:sfc:0-24hr diff:NAve=0 + wgrib rmf_gra_2016040600.24 | grep APCP + 291:28360360:d=16040600:APCP:kpds5=61:kpds6=1:kpds7=0: \ + TR=5:P1=0:P2=24:TimeU=1:sfc:0-24hr diff:NAve=0 - pcp_combine is looking in "rmf_gra_2016040600.24" for a 24 hour - *accumulation*, but since the time range indicator is no 4, it doesn't - find a match. + pcp_combine is looking in "rmf_gra_2016040600.24" for a 24 hour + *accumulation*, but since the time range indicator is not 4, it doesn't + find a match. - If possible switch the time range indicator to 4 on the GRIB files. If - this is not possible, there is another workaround. Instead of telling - pcp_combine to look for a particular accumulation interval, give it a - more complete description of the chosen field to use from each file. - Here is an example: + If possible switch the time range indicator to 4 on the GRIB files. If + this is not possible, there is another workaround. Instead of telling + pcp_combine to look for a particular accumulation interval, give it a + more complete description of the chosen field to use from each file. + Here is an example: - .. code-block:: none + .. code-block:: none - pcp_combine -add rmf_gra_2016040600.24 'name="APCP"; level="L0-24";' \ - rmf_gra_2016040600_APCP_00_24.nc + pcp_combine -add rmf_gra_2016040600.24 'name="APCP"; level="L0-24";' \ + rmf_gra_2016040600_APCP_00_24.nc - The resulting file should have the accumulation listed at - 24h rather than 0-24. + The resulting file should have the accumulation listed at + 24h rather than 0-24. Q. How do I use Pcp-Combine as a pass-through to simply reformat from GRIB to NetCDF or to change output variable name? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer - - The pcp_combine tool is typically used to modify the accumulation interval - of precipitation amounts in model and/or analysis datasets. For example, - when verifying model output in GRIB format containing runtime accumulations - of precipitation, run the pcp_combine -subtract option every 6 hours to - create 6-hourly precipitation amounts. In this example, it is not really - necessary to run pcp_combine on the 6-hour GRIB forecast file since the - model output already contains the 0 to 6 hour accumulation. However, the - output of pcp_combine is typically passed to point_stat, grid_stat, or mode - for verification. Having the 6-hour forecast in GRIB format and all other - forecast hours in NetCDF format (output of pcp_combine) makes the logic - for configuring the other MET tools messy. To make the configuration - consistent for all forecast hours, one option is to choose to run - pcp_combine as a pass-through to simply reformat from GRIB to NetCDF. - Listed below is an example of passing a single record to the - pcp_combine -add option to do the reformatting: - - .. code-block:: none - - $MET_BUILD/bin/pcp_combine -add forecast_F06.grb \ - 'name="APCP"; level="A6";' \ - forecast_APCP_06_F06.nc -name APCP_06 - - Reformatting from GRIB to NetCDF may be done for any other reason the - user may have. For example, the -name option can be used to define the - NetCDF output variable name. Presuming this file is then passed to - another MET tool, the new variable name (CompositeReflectivity) will - appear in the output of downstream tools: - - .. code-block:: none - - $MET_BUILD/bin/pcp_combine -add forecast.grb \ - 'name="REFC"; level="L0"; GRIB1_ptv=129; lead_time="120000";' \ - forecast.nc -name CompositeReflectivity - -Q. How do I use “-pcprx" to run a project faster? +.. dropdown:: Answer + + The pcp_combine tool is typically used to modify the accumulation interval + of precipitation amounts in model and/or analysis datasets. For example, + when verifying model output in GRIB format containing runtime accumulations + of precipitation, run the pcp_combine -subtract option every 6 hours to + create 6-hourly precipitation amounts. In this example, it is not really + necessary to run pcp_combine on the 6-hour GRIB forecast file since the + model output already contains the 0 to 6 hour accumulation. However, the + output of pcp_combine is typically passed to point_stat, grid_stat, or mode + for verification. Having the 6-hour forecast in GRIB format and all other + forecast hours in NetCDF format (output of pcp_combine) makes the logic + for configuring the other MET tools messy. To make the configuration + consistent for all forecast hours, one option is to choose to run + pcp_combine as a pass-through to simply reformat from GRIB to NetCDF. + Listed below is an example of passing a single record to the + pcp_combine -add option to do the reformatting: + + .. code-block:: none + + $MET_BUILD/bin/pcp_combine -add forecast_F06.grb \ + 'name="APCP"; level="A6";' \ + forecast_APCP_06_F06.nc -name APCP_06 + + Reformatting from GRIB to NetCDF may be done for any other reason the + user may have. For example, the -name option can be used to define the + NetCDF output variable name. Presuming this file is then passed to + another MET tool, the new variable name (CompositeReflectivity) will + appear in the output of downstream tools: + + .. code-block:: none + + $MET_BUILD/bin/pcp_combine -add forecast.grb \ + 'name="REFC"; level="L0"; GRIB1_ptv=129; lead_time="120000";' \ + forecast.nc -name CompositeReflectivity + +Q. How do I use "-pcprx" to run a project faster? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - To run a project faster, the “-pcprx” option may be used to narrow the - search down to whatever regular expression you provide. Here are a two - examples: + To run a project faster, the “-pcprx” option may be used to narrow the + search down to whatever regular expression you provide. Here are two + examples: - .. code-block:: none + .. code-block:: none - # Only using Stage IV data (ST4) - pcp_combine -sum 00000000_000000 06 \ - 20161015_18 12 ST4.2016101518.APCP_12_SUM.nc -pcprx "ST4.*.06h" + # Only using Stage IV data (ST4) + pcp_combine -sum 00000000_000000 06 \ + 20161015_18 12 ST4.2016101518.APCP_12_SUM.nc -pcprx "ST4.*.06h" - # Specify that files starting with pgbq[number][number]be used: - pcp_combine \ - -sum 20160221_18 06 20160222_18 24 \ - gfs_APCP_24_20160221_18_F00_F24.nc \ - -pcpdir /scratch4/BMC/shout/ptmp/Andrew.Kren/pre2016c3_corr/temp \ - -pcprx 'pgbq[0-9][0-9].gfs.2016022118' -v 3 + # Specify that files starting with pgbq[number][number] be used: + pcp_combine \ + -sum 20160221_18 06 20160222_18 24 \ + gfs_APCP_24_20160221_18_F00_F24.nc \ + -pcpdir /scratch4/BMC/shout/ptmp/Andrew.Kren/pre2016c3_corr/temp \ + -pcprx 'pgbq[0-9][0-9].gfs.2016022118' -v 3 Q. How do I enter the time format correctly? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - Here is an **incorrect example** of running pcp_combine with sub-hourly - accumulation intervals: + Here is an **incorrect example** of running pcp_combine with sub-hourly + accumulation intervals: - .. code-block:: none + .. code-block:: none - # incorrect example: - pcp_combine -subtract forecast.grb 0055 \ - forecast2.grb 0005 forecast.nc -field APCP + # incorrect example: + pcp_combine -subtract forecast.grb 0055 \ + forecast2.grb 0005 forecast.nc -field APCP - The time signature is entered incorrectly. Let’s assume that "0055" - meant 0 hours and 55 minutes and "0005" meant 0 hours and 5 minutes. + The time signature is entered incorrectly. Let’s assume that "0055" + meant 0 hours and 55 minutes and "0005" meant 0 hours and 5 minutes. - Looking at the usage statement for pcp_combine (just type pcp_combine with - no arguments): "accum1" indicates the accumulation interval to be used - from in_file1 in HH[MMSS] format (required). + Looking at the usage statement for pcp_combine (just type pcp_combine with + no arguments): "accum1" indicates the accumulation interval to be used + from in_file1 in HH[MMSS] format (required). - The time format listed "HH[MMSS]" means specifying hours or - hours/minutes/seconds. The incorrect example is using hours/minutes. + The time format listed "HH[MMSS]" means specifying hours or + hours/minutes/seconds. The incorrect example is using hours/minutes. - Below is the **correct example**. Add the seconds to the end of the - time strings, like this: + Below is the **correct example**. Add the seconds to the end of the + time strings, like this: - .. code-block:: none + .. code-block:: none - # correct example: - pcp_combine -subtract forecast.grb 005500 \ - forecast2.grb 000500 forecast.nc -field APCP + # correct example: + pcp_combine -subtract forecast.grb 005500 \ + forecast2.grb 000500 forecast.nc -field APCP Q. How do I use Pcp-Combine when my GRIB data doesn't have the appropriate accumulation interval time range indicator? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - Run wgrib on the data files and the output is listed below: + Run wgrib on the data files and the output is listed below: - .. code-block:: none + .. code-block:: none - 279:503477484:d=15062313:APCP:kpds5=61:kpds6=1:kpds7=0:TR= 10:P1=3:P2=247:TimeU=0:sfc:1015min \ - fcst:NAve=0 \ - 279:507900854:d=15062313:APCP:kpds5=61:kpds6=1:kpds7=0:TR= 10:P1=3:P2=197:TimeU=0:sfc:965min \ - fcst:NAve=0 + 279:503477484:d=15062313:APCP:kpds5=61:kpds6=1:kpds7=0:TR= 10:P1=3:P2=247:TimeU=0:sfc:1015min \ + fcst:NAve=0 \ + 279:507900854:d=15062313:APCP:kpds5=61:kpds6=1:kpds7=0:TR= 10:P1=3:P2=197:TimeU=0:sfc:965min \ + fcst:NAve=0 - Notice the output which says "TR=10". TR means time range indicator and - a value of 10 means that the level information contains an instantaneous - forecast time, not an accumulation interval. + Notice the output which says "TR=10". TR means time range indicator and + a value of 10 means that the level information contains an instantaneous + forecast time, not an accumulation interval. - Here's a table describing the TR values: - http://www.nco.ncep.noaa.gov/pmb/docs/on388/table5.html + Here's a table describing the TR values: + http://www.nco.ncep.noaa.gov/pmb/docs/on388/table5.html - The default logic for pcp_combine is to look for GRIB code 61 (i.e. APCP) - defined with an accumulation interval (TR = 4). Since the data doesn't - meet that criteria, the default logic of pcp_combine won't work. The - arguments need to be more specific to tell pcp_combine exactly what to do. + The default logic for pcp_combine is to look for GRIB code 61 (i.e., APCP) + defined with an accumulation interval (TR = 4). Since the data doesn't + meet that criteria, the default logic of pcp_combine won't work. The + arguments need to be more specific to tell pcp_combine exactly what to do. - Try the command: + Try the command: - .. code-block:: none + .. code-block:: none - pcp_combine -subtract \ - forecast.grb 'name="APCP"; level="L0"; lead_time="165500";' \ - forecast2.grb 'name="APCP"; level="L0"; lead_time="160500";' \ - forecast.nc -name APCP_A005000 + pcp_combine -subtract \ + forecast.grb 'name="APCP"; level="L0"; lead_time="165500";' \ + forecast2.grb 'name="APCP"; level="L0"; lead_time="160500";' \ + forecast.nc -name APCP_A005000 - Some things to point out here: + Some things to point out here: - 1. Notice in the wgrib output that the forecast times are 1015 min and - 965 min. In HHMMSS format, that's "165500" and "160500". + 1. Notice in the wgrib output that the forecast times are 1015 min and + 965 min. In HHMMSS format, that's "165500" and "160500". - 2. An accumulation interval can’t be specified since the data - isn't stored that way. Instead, use a config file string to - describe the data to use. + 2. An accumulation interval can’t be specified since the data + isn't stored that way. Instead, use a config file string to + describe the data to use. - 3. The config file string specifies a "name" (APCP) and "level" string. - APCP - is defined at the surface, so a level value of 0 (L0) was specified. + 3. The config file string specifies a "name" (APCP) and "level" string. + APCP + is defined at the surface, so a level value of 0 (L0) was specified. - 4. Technically, the "lead_time" doesn’t need to be specified at all, - pcp_combine - would find the single APCP record in each input GRIB file and use them. - But just in case, the lead_time option was included to be extra - certain to get exactly the data that is needed. + 4. Technically, the "lead_time" doesn’t need to be specified at all, + pcp_combine + would find the single APCP record in each input GRIB file and use them. + But just in case, the lead_time option was included to be extra + certain to get exactly the data that is needed. - 5. The default output variable name pcp_combine would write would be - "APCP_L0". However, to indicate that its a 50-minute - "accumulation interval" use a - different output variable name (APCP_A005000). Any string name is - possible. Maybe "Precip50Minutes" or "RAIN50". But whatever string is - chosen will be used in the Grid-Stat, Point-Stat, or MODE config file - to tell that tool what variable to process. + 5. The default output variable name pcp_combine would write would be + "APCP_L0". However, to indicate that it's a 50-minute + "accumulation interval" use a + different output variable name (APCP_A005000). Any string name is + possible. Maybe "Precip50Minutes" or "RAIN50". But whatever string is + chosen will be used in the Grid-Stat, Point-Stat, or MODE config file + to tell that tool what variable to process. -Q. How do I use “-sum”, “-add”, and “-subtract“ to achieve the same accumulation interval? +Q. How do I use "-sum", "-add", and "-subtract" to achieve the same accumulation interval? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer - - Here is an example of using pcp_combine to put GFS into 24- hour intervals - for comparison against 24-hourly StageIV precipitation with GFS data - through the pcp_combine tool. Be aware that the 24-hour StageIV data is - defined as an accumulation from 12Z on one day to 12Z on the next day: - https://water.noaa.gov/about/precipitation-data-access - - Therefore, only the 24-hour StageIV data can be used to evaluate 12Z to - 12Z accumulations from the model. Alternatively, the 6- hour StageIV - accumulations could be used to evaluate any 24 hour accumulation from - the model. For the latter, run the 6-hour StageIV files through - pcp_combine to generate the desired 24-hour accumulation. - - Here is an example. Run pcp_combine to compute 24-hour accumulations for - GFS. In this example, process the 20150220 00Z initialization of GFS. - - .. code-block:: none - - pcp_combine \ - -sum 20150220_00 06 20150221_00 24 \ - gfs_APCP_24_20150220_00_F00_F24.nc \ - -pcprx "gfs_4_20150220_00.*grb2" \ - -pcpdir /d1/model_data/20150220 - - pcp_combine is looking in the */d1/SBU/GFS/model_data/20150220* directory - at files which match this regular expression "gfs_4_20150220_00.*grb2". - That directory contains data for 00, 06, 12, and 18 hour initializations, - but the "-pcprx" option narrows the search down to the 00 hour - initialization which makes it run faster. It inspects all the matching - files, looking for 6-hour APCP data to sum up to a 24-hour accumulation - valid at 20150221_00. This results in a 24-hour accumulation between - forecast hours 0 and 24. - - The following command will compute the 24-hour accumulation between - forecast hours 12 and 36: - - .. code-block:: none - - pcp_combine \ - -sum 20150220_00 06 20150221_12 24 \ - gfs_APCP_24_20150220_00_F12_F36.nc \ - -pcprx "gfs_4_20150220_00.*grb2" \ - -pcpdir /d1/model_data/20150220 - - The "-sum" command is meant to make things easier by searching the - directory. But instead of using "-sum", another option would be the - "- add" command. Explicitly list the 4 files that need to be extracted - from the 6-hour APCP and add them up to 24. In the directory structure, - the previous "-sum" job could be rewritten with "-add" like this: - - .. code-block:: none - - pcp_combine -add \ - /d1/model_data/20150220/gfs_4_20150220_0000_018.grb2 06 \ - /d1/model_data/20150220/gfs_4_20150220_0000_024.grb2 06 \ - /d1/model_data/20150220/gfs_4_20150220_0000_030.grb2 06 \ - /d1/model_data/20150220/gfs_4_20150220_0000_036.grb2 06 \ - gfs_APCP_24_20150220_00_F12_F36_add_option.nc - - This example explicitly tells pcp_combine which files to read and - what accumulation interval (6 hours) to extract from them. The resulting - output should be identical to the output of the "-sum" command. - -Q. What is the difference between “-sum” vs. “-add”? +.. dropdown:: Answer + + Here is an example of using pcp_combine to put GFS into 24-hour intervals + for comparison against 24-hourly StageIV precipitation with GFS data + through the pcp_combine tool. Be aware that the 24-hour StageIV data is + defined as an accumulation from 12Z on one day to 12Z on the next day: + https://water.noaa.gov/about/precipitation-data-access + + Therefore, only the 24-hour StageIV data can be used to evaluate 12Z to + 12Z accumulations from the model. Alternatively, the 6-hour StageIV + accumulations could be used to evaluate any 24 hour accumulation from + the model. For the latter, run the 6-hour StageIV files through + pcp_combine to generate the desired 24-hour accumulation. + + Here is an example. Run pcp_combine to compute 24-hour accumulations for + GFS. In this example, process the 20150220 00Z initialization of GFS. + + .. code-block:: none + + pcp_combine \ + -sum 20150220_00 06 20150221_00 24 \ + gfs_APCP_24_20150220_00_F00_F24.nc \ + -pcprx "gfs_4_20150220_00.*grb2" \ + -pcpdir /d1/model_data/20150220 + + pcp_combine is looking in the */d1/model_data/20150220* directory + at files which match this regular expression "gfs_4_20150220_00.*grb2". + That directory contains data for 00, 06, 12, and 18 hour initializations, + but the "-pcprx" option narrows the search down to the 00 hour + initialization which makes it run faster. It inspects all the matching + files, looking for 6-hour APCP data to sum up to a 24-hour accumulation + valid at 20150221_00. This results in a 24-hour accumulation between + forecast hours 0 and 24. + + The following command will compute the 24-hour accumulation between + forecast hours 12 and 36: + + .. code-block:: none + + pcp_combine \ + -sum 20150220_00 06 20150221_12 24 \ + gfs_APCP_24_20150220_00_F12_F36.nc \ + -pcprx "gfs_4_20150220_00.*grb2" \ + -pcpdir /d1/model_data/20150220 + + The "-sum" command is meant to make things easier by searching the + directory. But instead of using "-sum", another option would be the + "-add" command. Explicitly list the 4 files that need to be extracted + from the 6-hour APCP and add them up to 24. In the directory structure, + the previous "-sum" job could be rewritten with "-add" like this: + + .. code-block:: none + + pcp_combine -add \ + /d1/model_data/20150220/gfs_4_20150220_0000_018.grb2 06 \ + /d1/model_data/20150220/gfs_4_20150220_0000_024.grb2 06 \ + /d1/model_data/20150220/gfs_4_20150220_0000_030.grb2 06 \ + /d1/model_data/20150220/gfs_4_20150220_0000_036.grb2 06 \ + gfs_APCP_24_20150220_00_F12_F36_add_option.nc + + This example explicitly tells pcp_combine which files to read and + what accumulation interval (6 hours) to extract from them. The resulting + output should be identical to the output of the "-sum" command. + +Q. What is the difference between "-sum" vs. "-add"? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - The -sum and -add options both do the same thing. It's just that - '-sum' could find files more quickly with the use of the -pcprx flag. - This could also be accomplished by using a calling script. + The -sum and -add options both do the same thing. It's just that + '-sum' could find files more quickly with the use of the -pcprx flag. + This could also be accomplished by using a calling script. Q. How do I select a specific GRIB record? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - In this example, record 735 needs to be selected. + In this example, record 735 needs to be selected. - .. code-block:: none + .. code-block:: none - pcp_combine -add 20160101_i12_f015_HRRR_wrfnat.grb2 \ - 'name="APCP"; level="R735";' \ - -name "APCP_01" HRRR_wrfnat.20160101_i12_f015.nc + pcp_combine -add 20160101_i12_f015_HRRR_wrfnat.grb2 \ + 'name="APCP"; level="R735";' \ + -name "APCP_01" HRRR_wrfnat.20160101_i12_f015.nc - Instead of having the level as "L0", tell it to use "R735" to select - grib record 735. + Instead of having the level as "L0", tell it to use "R735" to select + GRIB record 735. Plot-Data-Plane --------------- @@ -1011,117 +1012,113 @@ Plot-Data-Plane Q. How do I inspect Gen-Vx-Mask output? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - Check to see if the call to Gen-Vx-Mask actually did create good output - with Plot-Data-Plane. The following commands assume that the MET - executables are found in your path. + Check to see if the call to Gen-Vx-Mask actually did create good output + with Plot-Data-Plane. The following commands assume that the MET + executables are found in your path. - .. code-block:: none + .. code-block:: none - plot_data_plane \ - out/gen_vx_mask/CONUS_poly.nc \ - out/gen_vx_mask/CONUS_poly.ps \ - 'name="CONUS"; level="(*,*)";' + plot_data_plane \ + out/gen_vx_mask/CONUS_poly.nc \ + out/gen_vx_mask/CONUS_poly.ps \ + 'name="CONUS"; level="(*,*)";' - View that postscript output file, using something like "gv" - for ghostview: + View that PostScript output file, using something like "gv" + for ghostview: - .. code-block:: none + .. code-block:: none - gv out/gen_vx_mask/CONUS_poly.ps + gv out/gen_vx_mask/CONUS_poly.ps - Please review a map of 0's and 1's over the USA to determine if the output - file is what the user expects. It always a good idea to start with - plot_data_plane when working with data to make sure MET - is plotting the data correctly and in the expected location. + Please review a map of 0's and 1's over the USA to determine if the output + file is what the user expects. It is always a good idea to start with + plot_data_plane when working with data to make sure MET + is plotting the data correctly and in the expected location. Q. How do I specify the GRIB version? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer - - When MET reads Gridded data files, it must determine the type of - file it's reading. The first thing it checks is the suffix of the file. - The following are all interpreted as GRIB1: .grib, .grb, and .gb. - While these mean GRIB2: .grib2, .grb2, and .gb2. +.. dropdown:: Answer - There are 2 choices to control how MET interprets a grib file. Renaming - the files to use a particular suffix, or keep them - named and explicitly tell MET to interpret them as GRIB1 or GRIB2 using - the "file_type" configuration option. + When MET reads gridded data files, it must determine the type of + file it's reading. The first thing it checks is the suffix of the file. + The following are all interpreted as GRIB1: .grib, .grb, and .gb. + While these mean GRIB2: .grib2, .grb2, and .gb2. - The examples below use the plot_data_plane tool to plot the data. Set + There are 2 choices to control how MET interprets a GRIB file. Renaming + the files to use a particular suffix, or keep them + named and explicitly tell MET to interpret them as GRIB1 or GRIB2 using + the "file_type" configuration option. - .. code-block:: none + The example below uses the plot_data_plane tool to plot the data. - "file_type = GRIB2;" + To keep the files named as they are, add "file_type = GRIB2;" + to all the MET configuration files (i.e., Grid-Stat, MODE, and so on) + that you use: - To keep the files named this as they are, add "file_type = GRIB2;" - to all the MET configuration files (i.e. Grid-Stat, MODE, and so on) - that you use: - - .. code-block:: none + .. code-block:: none - plot_data_plane \ - test_2.5_prog.grib \ - test_2.5_prog.ps \ - 'name="TSTM"; level="A0"; file_type=GRIB2;' \ - -plot_range 0 100 + plot_data_plane \ + test_2.5_prog.grib \ + test_2.5_prog.ps \ + 'name="TSTM"; level="A0"; file_type=GRIB2;' \ + -plot_range 0 100 Q. How do I test the variable naming convention? (Record number example.) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - Make sure MET can read GRIB2 data. Plot the data from that GRIB2 file - by running: + Make sure MET can read GRIB2 data. Plot the data from that GRIB2 file + by running: - .. code-block:: none + .. code-block:: none - plot_data_plane LTIA98_KWBR_201305180600.grb2 tmp_z2.ps 'name="TMP"; level="R2"; + plot_data_plane LTIA98_KWBR_201305180600.grb2 tmp_z2.ps 'name="TMP"; level="R2";' - "R2" tells MET to plot record number 2. Record numbers 1 and 2 both - contain temperature data and 2-meters. Here's some wgrib2 output: + "R2" tells MET to plot record number 2. Record numbers 1 and 2 both + contain temperature data at 2-meters. Here's some wgrib2 output: - .. code-block:: none + .. code-block:: none - 1:0:d=2013051806:TMP:2 m above ground:anl:analysis/forecast error 2:3323062:d=2013051806:TMP:2 m above ground:anl: + 1:0:d=2013051806:TMP:2 m above ground:anl:analysis/forecast error 2:3323062:d=2013051806:TMP:2 m above ground:anl: - The GRIB id info has been the same between records 1 and 2. + The GRIB id info has been the same between records 1 and 2. Q. How do I compute and verify wind speed? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - Here's how to compute and verify wind speed using MET. Good news, MET - already includes logic for deriving wind speed on the fly. The GRIB - abbreviation for wind speed is WIND. To request WIND from a GRIB1 or - GRIB2 file, MET first checks to see if it already exists in the current - file. If so, it'll use it as is. If not, it'll search for the corresponding - U and V records and derive wind speed to use on the fly. + Here's how to compute and verify wind speed using MET. Good news, MET + already includes logic for deriving wind speed on the fly. The GRIB + abbreviation for wind speed is WIND. To request WIND from a GRIB1 or + GRIB2 file, MET first checks to see if it already exists in the current + file. If so, it'll use it as is. If not, it'll search for the corresponding + U and V records and derive wind speed to use on the fly. - In this example the RTMA file is named rtma.grb2 and the UPP file is - named wrf.grb, please try running the following commands to - plot wind speed: + In this example the RTMA file is named rtma.grb2 and the UPP file is + named wrf.grb, please try running the following commands to + plot wind speed: - .. code-block:: none + .. code-block:: none - plot_data_plane wrf.grb wrf_wind.ps \ - 'name"WIND"; level="Z10";' -v 3 - plot_data_plane rtma.grb2 rtma_wind.ps \ - 'name"WIND"; level="Z10";' -v 3 + plot_data_plane wrf.grb wrf_wind.ps \ + 'name="WIND"; level="Z10";' -v 3 + plot_data_plane rtma.grb2 rtma_wind.ps \ + 'name="WIND"; level="Z10";' -v 3 - In the first call, the log message should be similar to this: + In the first call, the log message should be similar to this: - .. code-block:: none + .. code-block:: none - DEBUG 3: MetGrib1DataFile::data_plane_array() -> - Attempt to derive winds from U and V components. + DEBUG 3: MetGrib1DataFile::data_plane_array() -> + Attempt to derive winds from U and V components. - In the second one, this won't appear since wind speed already exists - in the RTMA file. + In the second one, this won't appear since wind speed already exists + in the RTMA file. Stat-Analysis ------------- @@ -1129,196 +1126,197 @@ Stat-Analysis Q. How does '-aggregate_stat' work? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer - - In Stat-Analysis, there is a "-vx_mask" job filtering option. That option - reads the VX_MASK column from the input STAT lines and applies string - matching with the values in that column. Presumably, all of the MPR lines - will have the value of "FULL" in the VX_MASK column. - - Stat-Analysis has the ability to read MPR lines and recompute statistics - from them using the same library code that the other MET tools use. The - job command options which begin with "-out" are used to specify settings - to be applied to the output of that process. For example, - the "-fcst_thresh" - option filters strings from the input "FCST_THRESH" header column. The - "-out_fcst_thresh" option defines the threshold to be applied to the output - of Stat-Analysis. So reading MPR lines and applying a threshold to define - contingency table statistics (CTS) would be done using the - "-out_fcst_thresh" option. - - Stat-Analysis does have the ability to filter MPR lat/lon locations - using the "-mask_poly" option for a lat/lon polyline and the "-mask_grid" - option to define a retention grid. - - However, there is currently no "-mask_sid" option. - - With MET-5.2 and later versions, one option is to apply column string - matching using the "-column_str" option to define the list of station - ID's you would like to aggregate. That job would look something like this: - - .. code-block:: none - - stat_analysis -lookin path/to/mpr/directory \ - -job aggregate_stat -line_type MPR -out_line_type CNT \ - -column_str OBS_SID SID1,SID2,SID3,...,SIDN \ - -set_hdr VX_MASK SID_GROUP_NAME \ - -out_stat mpr_to_cnt.stat - - Where SID1...SIDN is a comma-separated list of the station id's in the - group. Notice that a value for the output VX_MASK column using the - "-set_hdr" option has been specified. Otherwise, this would show a list - of the unique values found in that column. Presumably, all the input - VX_MASK columns say "FULL" so that's what the output would say. Use - "-set_hdr" to explicitly set the output value. +.. dropdown:: Answer + + In Stat-Analysis, there is a "-vx_mask" job filtering option. That option + reads the VX_MASK column from the input STAT lines and applies string + matching with the values in that column. Presumably, all of the MPR lines + will have the value of "FULL" in the VX_MASK column. + + Stat-Analysis has the ability to read MPR lines and recompute statistics + from them using the same library code that the other MET tools use. The + job command options which begin with "-out" are used to specify settings + to be applied to the output of that process. For example, + the "-fcst_thresh" + option filters strings from the input "FCST_THRESH" header column. The + "-out_fcst_thresh" option defines the threshold to be applied to the output + of Stat-Analysis. So reading MPR lines and applying a threshold to define + contingency table statistics (CTS) would be done using the + "-out_fcst_thresh" option. + + Stat-Analysis does have the ability to filter MPR lat/lon locations + using the "-mask_poly" option for a lat/lon polyline and the "-mask_grid" + option to define a retention grid. + + The "-mask_sid" option filters MPR lines using a station ID masking + file or a comma-separated list of station IDs. + + Alternatively, one option is to apply column string + matching using the "-column_str" option to define the list of station + ID's you would like to aggregate. That job would look something like this: + + .. code-block:: none + + stat_analysis -lookin path/to/mpr/directory \ + -job aggregate_stat -line_type MPR -out_line_type CNT \ + -column_str OBS_SID SID1,SID2,SID3,...,SIDN \ + -set_hdr VX_MASK SID_GROUP_NAME \ + -out_stat mpr_to_cnt.stat + + Where SID1...SIDN is a comma-separated list of the station IDs in the + group. Notice that a value for the output VX_MASK column using the + "-set_hdr" option has been specified. Otherwise, this would show a list + of the unique values found in that column. Presumably, all the input + VX_MASK columns say "FULL" so that's what the output would say. Use + "-set_hdr" to explicitly set the output value. Q. What is the best way to average the FSS scores within several days or even several months using 'Aggregate to Average Scores'? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - Below is the best way to aggregate together the Neighborhood Continuous - (NBRCNT) lines across multiple days, specifically the fractions skill - score (FSS). The Stat-Analysis tool is designed to do this. This example - is for aggregating scores for the accumulated precipitation (APCP) field. + Below is the best way to aggregate together the Neighborhood Continuous + (NBRCNT) lines across multiple days, specifically the fractions skill + score (FSS). The Stat-Analysis tool is designed to do this. This example + is for aggregating scores for the accumulated precipitation (APCP) field. - Run the "aggregate" job type in stat_analysis to do this: + Run the "aggregate" job type in stat_analysis to do this: - .. code-block:: none + .. code-block:: none - stat_analysis -lookin directory/file*_nbrcnt.txt \ - -job aggregate -line_type NBRCNT -by FCST_VAR,FCST_LEAD,FCST_THRESH,INTERP_MTHD,INTERP_PNTS -out_stat agg_nbrcnt.txt + stat_analysis -lookin directory/file*_nbrcnt.txt \ + -job aggregate -line_type NBRCNT -by FCST_VAR,FCST_LEAD,FCST_THRESH,INTERP_MTHD,INTERP_PNTS -out_stat agg_nbrcnt.txt - This job reads all the files that are passed to it on the command line with - the "-lookin" option. List explicit filenames to read them directly. - Listing a top-level directory name will search that directory for files - ending in ".stat". + This job reads all the files that are passed to it on the command line with + the "-lookin" option. List explicit filenames to read them directly. + Listing a top-level directory name will search that directory for files + ending in ".stat". - In this case, the job running is to "aggregate" the "NBRCNT" line type. + In this case, the job running is to "aggregate" the "NBRCNT" line type. - In this case, the "-by" option is being used and lists several header - columns. Stat-Analysis will run this job separately for each unique - combination of those header column entries. + In this case, the "-by" option is being used and lists several header + columns. Stat-Analysis will run this job separately for each unique + combination of those header column entries. - The output is printed to the screen, or use the "-out_stat" option to - also write the aggregated output to a file named "agg_nbrcnt.txt". + The output is printed to the screen, or use the "-out_stat" option to + also write the aggregated output to a file named "agg_nbrcnt.txt". Q. How do I use '-by' to capture unique entries? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - Here is a stat-analysis job that could be used to run, read the - MPR lines, define the probabilistic forecast thresholds, define the - single observation threshold, and compute a PSTD output line. - Using "-by FCST_VAR" tells it to run the job separately for - each unique entry found in the FCST_VAR column. + Here is a stat-analysis job that could be used to run, read the + MPR lines, define the probabilistic forecast thresholds, define the + single observation threshold, and compute a PSTD output line. + Using "-by FCST_VAR" tells it to run the job separately for + each unique entry found in the FCST_VAR column. - .. code-block:: none + .. code-block:: none - stat_analysis \ - -lookin point_stat_model2_120000L_20160501_120000V.stat \ - -job aggregate_stat -line_type MPR -out_line_type PSTD \ - -out_fcst_thresh ge0,ge0.1,ge0.2,ge0.3,ge0.4,ge0.5,ge0.6,ge0.7,ge0.8,ge0.9,ge1.0 \ - -out_obs_thresh eq1.0 \ - -by FCST_VAR \ - -out_stat out_pstd.txt + stat_analysis \ + -lookin point_stat_model2_120000L_20160501_120000V.stat \ + -job aggregate_stat -line_type MPR -out_line_type PSTD \ + -out_fcst_thresh ge0,ge0.1,ge0.2,ge0.3,ge0.4,ge0.5,ge0.6,ge0.7,ge0.8,ge0.9,ge1.0 \ + -out_obs_thresh eq1.0 \ + -by FCST_VAR \ + -out_stat out_pstd.txt - The output statistics are written to "out_pstd.txt". + The output statistics are written to "out_pstd.txt". Q. How do I use '-filter' to refine my output? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer - - Here is an example of running a Stat-Analysis filter job to discard any - CNT lines (continuous statistics) where the forecast rate and observation - rate are less than 0.05. This is an alternative way of tossing out those - cases without having to modify the source code. - - .. code-block:: none - - stat_analysis \ - -lookin out/grid_stat/grid_stat_120000L_20050807_120000V.stat \ - -job filter -dump_row filter_cts.txt -line_type CTS \ - -column_min BASER 0.05 -column_min FMEAN 0.05 - DEBUG 2: STAT Lines read = 436 - DEBUG 2: STAT Lines retained = 36 - DEBUG 2: - DEBUG 2: Processing Job 1: -job filter -line_type CTS -column_min BASER - 0.05 -column_min - FMEAN 0.05 -dump_row filter_cts.txt - DEBUG 1: Creating - STAT output file "filter_cts.txt" - FILTER: -job filter -line_type - CTS -column_min - BASER 0.05 -column_min - FMEAN 0.05 -dump_row filter_cts.txt - DEBUG 2: Job 1 used 36 out of 36 STAT lines. - - This job reads find 56 CTS lines, but only keeps 36 of them where both - the BASER and FMEAN columns are at least 0.05. +.. dropdown:: Answer + + Here is an example of running a Stat-Analysis filter job to discard any + CTS lines (contingency table statistics) where the forecast rate and observation + rate are less than 0.05. This is an alternative way of tossing out those + cases without having to modify the source code. + + .. code-block:: none + + stat_analysis \ + -lookin out/grid_stat/grid_stat_120000L_20050807_120000V.stat \ + -job filter -dump_row filter_cts.txt -line_type CTS \ + -column_min BASER 0.05 -column_min FMEAN 0.05 + DEBUG 2: STAT Lines read = 436 + DEBUG 2: STAT Lines retained = 36 + DEBUG 2: + DEBUG 2: Processing Job 1: -job filter -line_type CTS -column_min BASER + 0.05 -column_min + FMEAN 0.05 -dump_row filter_cts.txt + DEBUG 1: Creating + STAT output file "filter_cts.txt" + FILTER: -job filter -line_type + CTS -column_min + BASER 0.05 -column_min + FMEAN 0.05 -dump_row filter_cts.txt + DEBUG 2: Job 1 used 36 out of 36 STAT lines. + + This job reads 436 STAT lines, but only keeps the 36 CTS lines where both + the BASER and FMEAN columns are at least 0.05. Q. How do I use the “-by” flag to stratify results? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - Adding "-by FCST_VAR" is a great way to associate a single value, - of say RMSE, with each of the forecast variables (UGRD,VGRD and WIND). + Adding "-by FCST_VAR" is a great way to associate a single value, + of say RMSE, with each of the forecast variables (UGRD,VGRD and WIND). - Run the following job on the output from Grid-Stat generated when the - "make test" command is run: + Run the following job on the output from Grid-Stat generated when the + "make test" command is run: - .. code-block:: none + .. code-block:: none - stat_analysis -lookin out/grid_stat \ - -job aggregate_stat -line_type SL1L2 -out_line_type CNT \ - -by FCST_VAR,FCST_LEV \ - -out_stat cnt.txt + stat_analysis -lookin out/grid_stat \ + -job aggregate_stat -line_type SL1L2 -out_line_type CNT \ + -by FCST_VAR,FCST_LEV \ + -out_stat cnt.txt - The resulting cnt.txt file includes separate output for 6 different - FCST_VAR values at different levels. + The resulting cnt.txt file includes separate output for 6 different + FCST_VAR values at different levels. Q. How do I speed up run times? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - By default, Stat-Analysis has two options enabled which slow it down. - Disabling these two options will create quicker run times: + By default, Stat-Analysis has two options enabled which slow it down. + Disabling these two options will create quicker run times: - 1. The computation of rank correlation statistics, Spearman's Rank - Correlation and Kendall's Tau. Disable them using - "-rank_corr_flag FALSE". + 1. The computation of rank correlation statistics, Spearman's Rank + Correlation and Kendall's Tau. Disable them using + "-rank_corr_flag FALSE". - 2. The computation of bootstrap confidence intervals. Disable them using - "-n_boot_rep 0". + 2. The computation of bootstrap confidence intervals. Disable them using + "-n_boot_rep 0". - Two more suggestions for faster run times. + Two more suggestions for faster run times. - 1. Instead of using "-fcst_var u", use "-by fcst_var". This will compute - statistics separately for each unique entry found in the - FCST_VAR column. + 1. Instead of using "-fcst_var u", use "-by fcst_var". This will compute + statistics separately for each unique entry found in the + FCST_VAR column. - 2. Instead of using "-out" to write the output to a text file, - use "-out_stat" - which will write a full STAT output file, including all the - header columns. - This will create a long list of values in the OBTYPE column. - To avoid the - long, OBTYPE column value, manually set the output using - "-set_hdr OBTYPE ALL_TYPES". Or set its value to whatever is needed. + 2. Instead of using "-out" to write the output to a text file, + use "-out_stat" + which will write a full STAT output file, including all the + header columns. + This will create a long list of values in the OBTYPE column. + To avoid the + long, OBTYPE column value, manually set the output using + "-set_hdr OBTYPE ALL_TYPES". Or set its value to whatever is needed. - .. code-block:: none + .. code-block:: none - stat_analysis \ - -lookin diag_conv_anl.2015060100.stat \ - -job aggregate_stat -line_type MPR -out_line_type CNT -by FCST_VAR \ - -out_stat diag_conv_anl.2015060100_cnt.txt -set_hdr OBTYPE ALL_TYPES \ - -n_boot_rep 0 -rank_corr_flag FALSE -v 4 + stat_analysis \ + -lookin diag_conv_anl.2015060100.stat \ + -job aggregate_stat -line_type MPR -out_line_type CNT -by FCST_VAR \ + -out_stat diag_conv_anl.2015060100_cnt.txt -set_hdr OBTYPE ALL_TYPES \ + -n_boot_rep 0 -rank_corr_flag FALSE -v 4 - Adding the "-by FCST_VAR" option to compute stats for all variables and - runs quickly. + Adding the "-by FCST_VAR" option to compute stats for all variables and + runs quickly. TC-Stat ------- @@ -1326,67 +1324,67 @@ TC-Stat Q. How do I use the “-by” flag to stratify results? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - To perform tropical cyclone evaluations for multiple models use the - "-by AMODEL" option with the tc_stat tool. Here is an example. + To perform tropical cyclone evaluations for multiple models use the + "-by AMODEL" option with the tc_stat tool. Here is an example. - In this case the tc_stat job looked at the 48 hour lead time for the HWRF - and H3HW models. Without the “-by AMODEL” option, the output would be - all grouped together. + In this case the tc_stat job looked at the 48 hour lead time for the HWRF + and H3HW models. Without the “-by AMODEL” option, the output would be + all grouped together. - .. code-block:: none + .. code-block:: none - tc_stat \ - -lookin d2014_vx_20141117_reset/al/tc_pairs/tc_pairs_H3WI_* \ - -lookin d2014_vx_20141117_reset/al/tc_pairs/tc_pairs_HWFI_* \ - -job summary -lead 480000 -column TRACK -amodel HWFI,H3WI \ - -by AMODEL -out sample.out + tc_stat \ + -lookin d2014_vx_20141117_reset/al/tc_pairs/tc_pairs_H3WI_* \ + -lookin d2014_vx_20141117_reset/al/tc_pairs/tc_pairs_HWFI_* \ + -job summary -lead 480000 -column TRACK -amodel HWFI,H3WI \ + -by AMODEL -out sample.out - This will result in all 48 hour HWFI and H3WI track forecasts to be - aggregated (statistics and scores computed) for each model separately. + This will result in all 48 hour HWFI and H3WI track forecasts to be + aggregated (statistics and scores computed) for each model separately. Q. How do I use rapid intensification verification? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - To get the most output, run something like this: + To get the most output, run something like this: - .. code-block:: none + .. code-block:: none - tc_stat \ - -lookin path/to/tc_pairs/output \ - -job rirw -dump_row test \ - -out_line_type CTC,CTS,MPR + tc_stat \ + -lookin path/to/tc_pairs/output \ + -job rirw -dump_row test \ + -out_line_type CTC,CTS,MPR - By default, rapid intensification (RI) is defined as a 24-hour exact - change exceeding 30kts. To define RI differently, modify that definition - using the ADECK, BDECK, or both using -rirw_time, -rirw_exact, - and -rirw_thresh options. Set -rirw_window to something larger than 0 - to enable false alarms to be considered hits when they were "close enough" - in time. + By default, rapid intensification (RI) is defined as a 24-hour exact + change exceeding 30kts. To define RI differently, modify that definition + using the ADECK, BDECK, or both using -rirw_time, -rirw_exact, + and -rirw_thresh options. Set -rirw_window to something larger than 0 + to enable false alarms to be considered hits when they were "close enough" + in time. - .. code-block:: none + .. code-block:: none - tc_stat \ - -lookin path/to/tc_pairs/output \ - -job rirw -dump_row test \ - -rirw_time 36 -rirw_window 12 \ - -out_line_type CTC,CTS,MPR + tc_stat \ + -lookin path/to/tc_pairs/output \ + -job rirw -dump_row test \ + -rirw_time 36 -rirw_window 12 \ + -out_line_type CTC,CTS,MPR - To evaluate Rapid Weakening (RW) by setting "-rirw_thresh <=-30". - To stratify your results by lead time, you could add the - "-by LEAD" option. + To evaluate Rapid Weakening (RW) by setting "-rirw_thresh <=-30". + To stratify your results by lead time, you could add the + "-by LEAD" option. - .. code-block:: none + .. code-block:: none - tc_stat \ - -lookin path/to/tc_pairs/output \ - -job rirw -dump_row test \ - -rirw_time 36 -rirw_window 12 \ - -rirw_thresh <=-30 -by LEAD \ - -out_line_type CTC,CTS,MPR + tc_stat \ + -lookin path/to/tc_pairs/output \ + -job rirw -dump_row test \ + -rirw_time 36 -rirw_window 12 \ + -rirw_thresh <=-30 -by LEAD \ + -out_line_type CTC,CTS,MPR Utilities --------- @@ -1394,136 +1392,136 @@ Utilities Q. What would be an example of scripting to call MET? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - The following is an example of how to call MET from a bash script - including passing in variables. This shell script is listed below to run - Grid-Stat, call Plot-Data-Plane to plot the resulting difference field, - and call convert to reformat from PostScript to PNG. + The following is an example of how to call MET from a bash script + including passing in variables. This shell script is listed below to run + Grid-Stat, call Plot-Data-Plane to plot the resulting difference field, + and call convert to reformat from PostScript to PNG. - .. code-block:: none + .. code-block:: none - #!/bin/sh - for case in `echo "FCST OBS"`; do - export TO_GRID=${case} - grid_stat gfs.t00z.pgrb2.0p25.f000 \ - nam.t00z.conusnest.hiresf00.tm00.grib2 GridStatConfig - plot_data_plane \ - *TO_GRID_${case}*_pairs.nc TO_GRID_${case}.ps 'name="DIFF_TMP_P500_TMP_P500_FULL"; \ - level="(*,*)";' - convert -rotate 90 -background white -flatten TO_GRID_${case}.ps - TO_GRID_${case}.png - done + #!/bin/sh + for case in `echo "FCST OBS"`; do + export TO_GRID=${case} + grid_stat gfs.t00z.pgrb2.0p25.f000 \ + nam.t00z.conusnest.hiresf00.tm00.grib2 GridStatConfig + plot_data_plane \ + *TO_GRID_${case}*_pairs.nc TO_GRID_${case}.ps 'name="DIFF_TMP_P500_TMP_P500_FULL"; \ + level="(*,*)";' + convert -rotate 90 -background white -flatten TO_GRID_${case}.ps + TO_GRID_${case}.png + done Q. How do I convert TRMM data files? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer - - Here is an example of NetCDF that the MET software is not expecting. Here - is an option for accessing that same TRMM data, following links from the - MET website: - http://dtcenter.org/community-code/model-evaluation-tools-met/input-data - - .. code-block:: none - - # Pull binary 3-hourly TRMM data file - wget - ftp://disc2.nascom.nasa.gov/data/TRMM/Gridded/3B42_V7/201009/3B42.100921.00z.7. - precipitation.bin - # Pull Rscript from MET website - wget http://dtcenter.org/sites/default/files/community-code/met/r-scripts/trmmbin2nc.R - # Edit that Rscript by setting - out_lat_ll = -50 - out_lon_ll = 0 - out_lat_ur = 50 - out_lon_ur = 359.75 - # Run the Rscript - Rscript trmmbin2nc.R 3B42.100921.00z.7.precipitation.bin \ - 3B42.100921.00z.7.precipitation.nc - # Plot the result - plot_data_plane 3B42.100921.00z.7.precipitation.nc \ - 3B42.100921.00z.7.precipitation.ps 'name="APCP_03"; level="(*,*)";' - - It may be possible that the domain of the data is smaller. - Here are some options: - - 1. In that Rscript, choose different boundaries (i.e. out_lat/lon_ll/ur) - to specify the tile of data to be selected. - - 2. As of version 5.1, MET includes support for regridding the - data it reads. Keep TRMM on it's native domain and use the - MET tools to do the regridding. - For example, the Regrid-Data-Plane" tool reads a NetCDF file, regrids - the data, and writes a NetCDF file. Alternatively, the "regrid" section - of the configuration files for the MET tools may be used to do the - regridding on the fly. For example, run Grid-Stat to compare to - the model output to TRMM and say - - .. code-block:: none - - "regrid = { field = FCST; - ...}" - - That tells Grid-Stat to automatically regrid the TRMM observations to - the model domain. +.. dropdown:: Answer + + Here is an example of NetCDF that the MET software is not expecting. Here + is an option for accessing that same TRMM data, following links from the + MET website: + http://dtcenter.org/community-code/model-evaluation-tools-met/input-data + + .. code-block:: none + + # Pull binary 3-hourly TRMM data file + wget + ftp://disc2.nascom.nasa.gov/data/TRMM/Gridded/3B42_V7/201009/3B42.100921.00z.7. + precipitation.bin + # Pull Rscript from MET website + wget http://dtcenter.org/sites/default/files/community-code/met/r-scripts/trmmbin2nc.R + # Edit that Rscript by setting + out_lat_ll = -50 + out_lon_ll = 0 + out_lat_ur = 50 + out_lon_ur = 359.75 + # Run the Rscript + Rscript trmmbin2nc.R 3B42.100921.00z.7.precipitation.bin \ + 3B42.100921.00z.7.precipitation.nc + # Plot the result + plot_data_plane 3B42.100921.00z.7.precipitation.nc \ + 3B42.100921.00z.7.precipitation.ps 'name="APCP_03"; level="(*,*)";' + + It may be possible that the domain of the data is smaller. + Here are some options: + + 1. In that Rscript, choose different boundaries (i.e., out_lat/lon_ll/ur) + to specify the tile of data to be selected. + + 2. As of version 5.1, MET includes support for regridding the + data it reads. Keep TRMM on its native domain and use the + MET tools to do the regridding. + For example, the Regrid-Data-Plane tool reads a NetCDF file, regrids + the data, and writes a NetCDF file. Alternatively, the "regrid" section + of the configuration files for the MET tools may be used to do the + regridding on the fly. For example, run Grid-Stat to compare to + the model output to TRMM and say + + .. code-block:: none + + "regrid = { field = FCST; + ...}" + + That tells Grid-Stat to automatically regrid the TRMM observations to + the model domain. Q. How do I convert a PostScript to png? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - Use the linux “convert” tool to convert a Plot-Data-Plane PostScript - file to a png: + Use the linux “convert” tool to convert a Plot-Data-Plane PostScript + file to a png: - .. code-block:: none + .. code-block:: none - convert -rotate 90 -background white plot_dbz.ps plot_dbz.png + convert -rotate 90 -background white plot_dbz.ps plot_dbz.png - To convert a MODE PostScript to png + To convert a MODE PostScript to png - .. code-block:: none + .. code-block:: none - convert mode_out.ps mode_out.png + convert mode_out.ps mode_out.png - Will result in all 6-7 pages in the PostScript file be written out to a - seperate .png with the following naming convention: + Will result in all 6-7 pages in the PostScript file be written out to a + separate .png with the following naming convention: - mode_out-0.png, mode_out-1.png, mode_out-2.png, etc. + mode_out-0.png, mode_out-1.png, mode_out-2.png, etc. Q. How does pairwise differences using plot_tcmpr.R work? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - One necessary step in computing pairwise differences is "event equalizing" - the data. This means extracting a subset of cases that are common to - both models. + One necessary step in computing pairwise differences is "event equalizing" + the data. This means extracting a subset of cases that are common to + both models. - While the tc_stat tool does not compute pairwise differences, it can apply - the "event_equalization" logic to extract the cases common to two models. - This is done using the config file "event_equal = TRUE;" option or - setting "-event_equal true" on the command line. + While the tc_stat tool does not compute pairwise differences, it can apply + the "event_equalization" logic to extract the cases common to two models. + This is done using the config file "event_equal = TRUE;" option or + setting "-event_equal true" on the command line. - Most of the hurricane track analysis and plotting is done using the - plot_tcmpr.R Rscript. It makes a call to the tc_stat tool to track - data down to the desired subset, compute pairwise differences if needed, - and then plot the result. + Most of the hurricane track analysis and plotting is done using the + plot_tcmpr.R Rscript. It makes a call to the tc_stat tool to track + data down to the desired subset, compute pairwise differences if needed, + and then plot the result. - .. code-block:: none + .. code-block:: none - Rscript ${MET_BASE}/Rscripts/plot_tcmpr.R \ - -lookin tc_pairs_output.tcst \ - -filter '-amodel AHWI,GFSI' \ - -series AMODEL AHWI,GFSI,AHWI-GFSI \ - -plot MEAN,BOXPLOT + Rscript ${MET_BASE}/Rscripts/plot_tcmpr.R \ + -lookin tc_pairs_output.tcst \ + -filter '-amodel AHWI,GFSI' \ + -series AMODEL AHWI,GFSI,AHWI-GFSI \ + -plot MEAN,BOXPLOT - The resulting plots include three series - one for AHWI, one for GFSI, - and one for their pairwise difference. + The resulting plots include three series - one for AHWI, one for GFSI, + and one for their pairwise difference. - It's a bit cumbersome to understand all the options available, but this may - be really useful. If nothing else, it could be adapted to dump out the - pairwise differences that are needed. + It's a bit cumbersome to understand all the options available, but this may + be really useful. If nothing else, it could be adapted to dump out the + pairwise differences that are needed. Miscellaneous @@ -1532,282 +1530,282 @@ Miscellaneous Q. Regrid-Data-Plane - How do I define a LatLon grid? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - Here is an example of the NetCDF variable attributes that MET uses to - define a LatLon grid: + Here is an example of the NetCDF variable attributes that MET uses to + define a LatLon grid: - .. code-block:: none + .. code-block:: none - :Projection = "LatLon" ; - :lat_ll = "25.063000 degrees_north" ; - :lon_ll = "-124.938000 degrees_east" ; - :delta_lat = "0.125000 degrees" ; - :delta_lon = "0.125000 degrees" ; - :Nlat = "224 grid_points" ; - :Nlon = "464 grid_points" ; + :Projection = "LatLon" ; + :lat_ll = "25.063000 degrees_north" ; + :lon_ll = "-124.938000 degrees_east" ; + :delta_lat = "0.125000 degrees" ; + :delta_lon = "0.125000 degrees" ; + :Nlat = "224 grid_points" ; + :Nlon = "464 grid_points" ; - This can be created by running the Regrid-Data-Plane" tool to regrid - some GFS data to a LatLon grid: + This can be created by running the Regrid-Data-Plane tool to regrid + some GFS data to a LatLon grid: - .. code-block:: none + .. code-block:: none - regrid_data_plane \ - gfs_2012040900_F012.grib G110 \ - gfs_g110.nc -field 'name="TMP"; level="Z2";' + regrid_data_plane \ + gfs_2012040900_F012.grib G110 \ + gfs_g110.nc -field 'name="TMP"; level="Z2";' - Use ncdump to look at the attributes. As an exercise, try defining - these global attributes (and removing the other projection-related ones) - and then try again. + Use ncdump to look at the attributes. As an exercise, try defining + these global attributes (and removing the other projection-related ones) + and then try again. Q. Pre-processing - How do I use wgrib2, pcp_combine regrid and reformat to format NetCDF files? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - If you are extracting only one or two fields from a file, using MET's - Regrid-Data-Plane can be used to generate a Lat-Lon projection. If - regridding all fields, the wgrib2 utility may be more useful. Here's an - example of using wgrib2 and pcp_combine to generate NetCDF files - MET can read: + If you are extracting only one or two fields from a file, using MET's + Regrid-Data-Plane can be used to generate a Lat-Lon projection. If + regridding all fields, the wgrib2 utility may be more useful. Here's an + example of using wgrib2 and pcp_combine to generate NetCDF files + MET can read: - .. code-block:: none + .. code-block:: none - wgrib2 gfsrain06.grb -new_grid latlon 112:131:0.1 \ - 25:121:0.1 gfsrain06_regrid.grb2 + wgrib2 gfsrain06.grb -new_grid latlon 112:131:0.1 \ + 25:121:0.1 gfsrain06_regrid.grb2 - And then run that GRIB2 file through pcp_combine using the "-add" option - with only one file provided: + And then run that GRIB2 file through pcp_combine using the "-add" option + with only one file provided: - .. code-block:: none + .. code-block:: none - pcp_combine -add gfsrain06_regrid.grb2 'name="APCP"; \ - level="A6";' gfsrain06_regrid.nc + pcp_combine -add gfsrain06_regrid.grb2 'name="APCP"; \ + level="A6";' gfsrain06_regrid.nc - Then the output NetCDF file does not have this problem: + Then the output NetCDF file does not have this problem: - .. code-block:: none + .. code-block:: none - ncdump -h 2a_wgrib2_regrid.nc | grep "_ll" - :lat_ll = "25.000000 degrees_north" ; - :lon_ll = "112.000000 degrees_east" ; + ncdump -h 2a_wgrib2_regrid.nc | grep "_ll" + :lat_ll = "25.000000 degrees_north" ; + :lon_ll = "112.000000 degrees_east" ; Q. TC-Pairs - How do I get rid of WARNING: TrackInfo Using Specify Model Suffix? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer - - Below is a command example to run: - - .. code-block:: none - - tc_pairs \ - -adeck aep142014.h4hw.dat \ - -bdeck bep142014.dat \ - -config TCPairsConfig_v5.0 \ - -out tc_pairs_v5.0_patch \ - -log tc_pairs_v5.0_patch.log \ - -v 3 - - Below is a warning message: - - .. code-block:: none - - WARNING: TrackInfo::add(const ATCFLine &) -> - skipping ATCFLine since the valid time is not - increasing (20140801_000000 < 20140806_060000): - WARNING: AL, 03, 2014080100, 03, H4HW, 000, - 120N, 547W, 38, 1009, XX, 34, NEQ, 0084, 0000, - 0000, 0083, -99, -99, 59, 0, 0, , 0, , 0, 0, - - As a sanity check, the MET-TC code makes sure that the valid time of - the track data doesn't go backwards in time. This warning states that - this is - occurring. The very likely reason for this is that the data being used - are probably passing tc_pairs duplicate track data. - - Using grep, notice that the same track data shows up in - "aal032014.h4hw.dat" and "aal032014_hfip_d2014_BERTHA.dat". Try this: - - .. code-block:: none - - grep H4HW aal*.dat | grep 2014080100 | grep ", 000," - aal032014.h4hw.dat:AL, 03, 2014080100, 03, H4HW, 000, - 120N, 547W, 38, 1009, XX, 34, NEQ, 0084, - 0000, 0000, 0083, -99, -99, 59, 0, 0, , - 0, , 0, 0, , , , , 0, 0, 0, 0, THERMO PARAMS, - -9999, -9999, -9999, Y, 10, DT, -999 - aal032014_hfip_d2014_BERTHA.dat:AL, 03, 2014080100, - 03, H4HW, 000, 120N, 547W, 38, 1009, XX, 34, NEQ, - 0084, 0000, 0000, 0083, -99, -99, 59, 0, 0, , 0, , 0, - 0, , , , , 0, 0, 0, 0, THERMOPARAMS, -9999 ,-9999 , - -9999 ,Y ,10 ,DT ,-999 - - Those 2 lines are nearly identical, except for the spelling of - "THERMO PARAMS" with a space vs "THERMOPARAMS" with no space. - - Passing tc_pairs duplicate track data results in this sort of warning. - The DTC had the same sort of problem when setting up a real-time - verification system. The same track data was making its way into - multiple ATCF files. - - If this really is duplicate track data, work on the logic for where/how - to store the track data. However, if the H4HW data in the first file - actually differs from that in the second file, there is another option. - You can specify a model suffix to be used for each ADECK source, as in - this example (suffix=_EXP): - - .. code-block:: none - - tc_pairs \ - -adeck aal032014.h4hw.dat suffix=_EXP \ - -adeck aal032014_hfip_d2014_BERTHA.dat \ - -bdeck bal032014.dat \ - -config TCPairsConfig_match \ - -out tc_pairs_v5.0_patch \ - -log tc_pairs_v5.0_patch.log -v 3 - - Any model names found in "aal032014.h4hw.dat" will now have _EXP tacked - onto the end. Note that if a list of model names in the TCPairsConfig file - needs specifying, include the _EXP variants to get them to show up in - the output or it won’t show up. - - That'll get rid of the warnings because they will be storing the track - data from the first source using a slightly different model name. This - feature was added for users who are testing multiple versions of a - model on the same set of storms. They might be using the same ATCF ID - in all their output. But this enables them to distinguish the output - in tc_pairs. +.. dropdown:: Answer + + Below is a command example to run: + + .. code-block:: none + + tc_pairs \ + -adeck aep142014.h4hw.dat \ + -bdeck bep142014.dat \ + -config TCPairsConfig_v5.0 \ + -out tc_pairs_v5.0_patch \ + -log tc_pairs_v5.0_patch.log \ + -v 3 + + Below is a warning message: + + .. code-block:: none + + WARNING: TrackInfo::add(const ATCFLine &) -> + skipping ATCFLine since the valid time is not + increasing (20140801_000000 < 20140806_060000): + WARNING: AL, 03, 2014080100, 03, H4HW, 000, + 120N, 547W, 38, 1009, XX, 34, NEQ, 0084, 0000, + 0000, 0083, -99, -99, 59, 0, 0, , 0, , 0, 0, + + As a sanity check, the MET-TC code makes sure that the valid time of + the track data doesn't go backwards in time. This warning states that + this is + occurring. The very likely reason for this is that the data being used + are probably passing tc_pairs duplicate track data. + + Using grep, notice that the same track data shows up in + "aal032014.h4hw.dat" and "aal032014_hfip_d2014_BERTHA.dat". Try this: + + .. code-block:: none + + grep H4HW aal*.dat | grep 2014080100 | grep ", 000," + aal032014.h4hw.dat:AL, 03, 2014080100, 03, H4HW, 000, + 120N, 547W, 38, 1009, XX, 34, NEQ, 0084, + 0000, 0000, 0083, -99, -99, 59, 0, 0, , + 0, , 0, 0, , , , , 0, 0, 0, 0, THERMO PARAMS, + -9999, -9999, -9999, Y, 10, DT, -999 + aal032014_hfip_d2014_BERTHA.dat:AL, 03, 2014080100, + 03, H4HW, 000, 120N, 547W, 38, 1009, XX, 34, NEQ, + 0084, 0000, 0000, 0083, -99, -99, 59, 0, 0, , 0, , 0, + 0, , , , , 0, 0, 0, 0, THERMOPARAMS, -9999 ,-9999 , + -9999 ,Y ,10 ,DT ,-999 + + Those 2 lines are nearly identical, except for the spelling of + "THERMO PARAMS" with a space vs "THERMOPARAMS" with no space. + + Passing tc_pairs duplicate track data results in this sort of warning. + The DTC had the same sort of problem when setting up a real-time + verification system. The same track data was making its way into + multiple ATCF files. + + If this really is duplicate track data, work on the logic for where/how + to store the track data. However, if the H4HW data in the first file + actually differs from that in the second file, there is another option. + You can specify a model suffix to be used for each ADECK source, as in + this example (suffix=_EXP): + + .. code-block:: none + + tc_pairs \ + -adeck aal032014.h4hw.dat suffix=_EXP \ + -adeck aal032014_hfip_d2014_BERTHA.dat \ + -bdeck bal032014.dat \ + -config TCPairsConfig_match \ + -out tc_pairs_v5.0_patch \ + -log tc_pairs_v5.0_patch.log -v 3 + + Any model names found in "aal032014.h4hw.dat" will now have _EXP tacked + onto the end. Note that if a list of model names in the TCPairsConfig file + needs specifying, include the _EXP variants to get them to show up in + the output or it won’t show up. + + That'll get rid of the warnings because they will be storing the track + data from the first source using a slightly different model name. This + feature was added for users who are testing multiple versions of a + model on the same set of storms. They might be using the same ATCF ID + in all their output. But this enables them to distinguish the output + in tc_pairs. Q. Why is the grid upside down? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer - - The user provides a gridded data file to MET and it runs without error, - but the data is packed upside down. - - Try using the "file_type" entry. The "file_type" entry specifies the - input file type (e.g. GRIB1, GRIB2, NETCDF_MET, NETCDF_WRF, NETCDF_PINT, NETCDF_NCCF) - rather than letting the code determine it itself. For valid file_type - values, see "File types" in the *data/config/ConfigConstants* file. This - entry should be defined within the "fcst" or "obs" dictionaries. - Sometimes, directly specifying the type of file will help MET figure - out what to properly do with the data. - - Another option is to use the Regrid-Data-Plane tool. The Regrid-Data-Plane - tool may be run to read data from any gridded data file MET supports - (i.e. GRIB1, GRIB2, and a variety of NetCDF formats), interpolate to a - user-specified grid, and write the field(s) out in NetCDF format. See the - Regrid-Data-Plane tool :numref:`regrid-data-plane` in the MET - User's Guide for more - detailed information. While the Regrid-Data-Plane tool is useful as a - stand-alone tool, the capability is also included to automatically regrid - data in most of the MET tools that handle gridded data. This "regrid" - entry is a dictionary containing information about how to handle input - gridded data files. The "regird" entry specifies regridding logic and - has a "to_grid" entry that can be set to NONE, FCST, OBS, a named grid, - the path to a gridded data file defining the grid, or an explicit grid - specification string. See the :ref:`regrid` entry in - the Configuration File Overview in the MET User's Guide for a more detailed - description of the configuration file entries that control automated - regridding. - - A single model level can be plotted using the plot_data_plane utility. - This tool can assist the user by showing the data to be verified to - ensure that times and locations matchup as expected. +.. dropdown:: Answer + + The user provides a gridded data file to MET and it runs without error, + but the data is packed upside down. + + Try using the "file_type" entry. The "file_type" entry specifies the + input file type (e.g., GRIB1, GRIB2, NETCDF_MET, NETCDF_WRF, NETCDF_PINT, NETCDF_NCCF) + rather than letting the code determine it itself. For valid file_type + values, see "File types" in the *data/config/ConfigConstants* file. This + entry should be defined within the "fcst" or "obs" dictionaries. + Sometimes, directly specifying the type of file will help MET figure + out what to properly do with the data. + + Another option is to use the Regrid-Data-Plane tool. The Regrid-Data-Plane + tool may be run to read data from any gridded data file MET supports + (i.e., GRIB1, GRIB2, and a variety of NetCDF formats), interpolate to a + user-specified grid, and write the field(s) out in NetCDF format. See the + Regrid-Data-Plane tool :numref:`regrid-data-plane` in the MET + User's Guide for more + detailed information. While the Regrid-Data-Plane tool is useful as a + stand-alone tool, the capability is also included to automatically regrid + data in most of the MET tools that handle gridded data. This "regrid" + entry is a dictionary containing information about how to handle input + gridded data files. The "regrid" entry specifies regridding logic and + has a "to_grid" entry that can be set to NONE, FCST, OBS, a named grid, + the path to a gridded data file defining the grid, or an explicit grid + specification string. See the :ref:`regrid` entry in + the Configuration File Overview in the MET User's Guide for a more detailed + description of the configuration file entries that control automated + regridding. + + A single model level can be plotted using the plot_data_plane utility. + This tool can assist the user by showing the data to be verified to + ensure that times and locations match up as expected. Q. Why was the MET written largely in C++ instead of FORTRAN? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - MET relies upon the object-oriented aspects of C++, particularly in - using the MODE tool. Due to time and budget constraints, it also makes - use of a pre-existing forecast verification library that was developed - at NCAR. + MET relies upon the object-oriented aspects of C++, particularly in + using the MODE tool. Due to time and budget constraints, it also makes + use of a pre-existing forecast verification library that was developed + at NCAR. Q. How does MET differ from the previously mentioned existing verification packages? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - MET is an actively maintained, evolving software package that is being - made freely available to the public through controlled version releases. + MET is an actively maintained, evolving software package that is being + made freely available to the public through controlled version releases. Q. Will the MET work on data in native model coordinates? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - No - it will not. In the future, we may add options to allow additional - model grid coordinate systems. + No - it will not. In the future, we may add options to allow additional + model grid coordinate systems. Q. How do I get help if my questions are not answered in the User's Guide? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - First, look on our - `MET User's Guide website `_. - If that doesn't answer your question, create a post in the - `METplus GitHub Discussions Forum `_. + First, look on our + `MET User's Guide website `_. + If that doesn't answer your question, create a post in the + `METplus GitHub Discussions Forum `_. Q. What graphical features does MET provide? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer - - MET provides some :ref:`plotting and graphics support`. - The plotting tools, including plot_point_obs and plot_data_plane, can help - users visualize the data. - - MET is intended to be a set of command line tools for evaluating forecast - quality. So, the development effort is focused on providing the latest, - state of the art verification approaches, rather than on providing nice - plotting features. However, the ASCII output statistics of MET may - be plotted - with a wide variety of plotting packages, including R, NCL, IDL, - and GNUPlot. - METViewer is also currently being developed and used by the DTC and NOAA - It creates basic plots of MET output verification statistics. The types of - plots include series plots with confidence intervals, box plots, - x-y scatter plots and histograms. - - R is a language and environment for statistical computing and graphics. - It's a free package that runs on most operating systems and provides nice - plotting features and a wide array of powerful statistical analysis tools. - There are sample scripts on the - `MET website `_ - that you can use and modify to perform the type of analysis you need. If - you create your own scripts, we encourage you to submit them to us - through the - `METplus GitHub Discussions Forum `_ - so that we can post them for other users. +.. dropdown:: Answer + + MET provides some :ref:`plotting and graphics support`. + The plotting tools, including plot_point_obs and plot_data_plane, can help + users visualize the data. + + MET is intended to be a set of command line tools for evaluating forecast + quality. So, the development effort is focused on providing the latest, + state of the art verification approaches, rather than on providing nice + plotting features. However, the ASCII output statistics of MET may + be plotted + with a wide variety of plotting packages, including R, NCL, IDL, + and GNUPlot. + METViewer is also currently being developed and used by the DTC and NOAA. + It creates basic plots of MET output verification statistics. The types of + plots include series plots with confidence intervals, box plots, + x-y scatter plots and histograms. + + R is a language and environment for statistical computing and graphics. + It's a free package that runs on most operating systems and provides nice + plotting features and a wide array of powerful statistical analysis tools. + There are sample scripts on the + `MET website `_ + that you can use and modify to perform the type of analysis you need. If + you create your own scripts, we encourage you to submit them to us + through the + `METplus GitHub Discussions Forum `_ + so that we can post them for other users. Q. How do I find the version of the tool I am using? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - Type the name of the tool followed by **--version**. For example, - type “pb2nc **--version**”. + Type the name of the tool followed by **--version**. For example, + type “pb2nc **--version**”. Q. What are MET's conventions for latitude, longitude, azimuth and bearing angles? ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ - .. dropdown:: Answer +.. dropdown:: Answer - MET considers north latitude and east longitude positive. However, - internally MET considers east longitude negative so users may encounter - DEBUG statements with longitude of a different sign than they provided - (e.g. for observation locations or grid metadata). Latitudes have - range from :math:`-90^\circ` to :math:`+90^\circ`. Longitudes have - range from :math:`-180^\circ` to :math:`+180^\circ`. Plane angles such - as azimuths and bearing (example: horizontal wind direction) have - range :math:`0^\circ` to :math:`360^\circ` and are measured clockwise - from the north. + MET considers north latitude and east longitude positive. However, + internally MET considers east longitude negative so users may encounter + DEBUG statements with longitude of a different sign than they provided + (e.g., for observation locations or grid metadata). Latitudes have + range from :math:`-90^\circ` to :math:`+90^\circ`. Longitudes have + range from :math:`-180^\circ` to :math:`+180^\circ`. Plane angles such + as azimuths and bearing (example: horizontal wind direction) have + range :math:`0^\circ` to :math:`360^\circ` and are measured clockwise + from the north. .. _Troubleshooting: @@ -1825,157 +1823,157 @@ on other things to check if you are having problems installing or running MET. MET Won't Compile ----------------- - .. dropdown:: Troubleshooting Help +.. dropdown:: Troubleshooting Help - * Have you specified the locations of NetCDF, GNU Scientific Library, - and BUFRLIB, and optional additional libraries using corresponding - MET\_ environment variables prior to running configure? + * Have you specified the locations of NetCDF, GNU Scientific Library, + and BUFRLIB, and optional additional libraries using corresponding + MET\_ environment variables prior to running configure? - * Have these libraries been compiled and installed using the same set - of compilers used to build MET? + * Have these libraries been compiled and installed using the same set + of compilers used to build MET? BUFRLIB Errors During MET Installation -------------------------------------- - .. dropdown:: Troubleshooting Help +.. dropdown:: Troubleshooting Help - .. code-block:: none + .. code-block:: none - error message: /usr/bin/ld: cannot find -lbufr - The linker can not find the BUFRLIB library archive file it needs. + error message: /usr/bin/ld: cannot find -lbufr + The linker can not find the BUFRLIB library archive file it needs. - export MET_BUFRLIB=/home/username/BUFRLIB_v11.3.0:$MET_BUFRLIB + export MET_BUFRLIB=/home/username/BUFRLIB_v11.3.0:$MET_BUFRLIB - It isn't making it's way into the configuration because BUFRLIB_v11.3.0 - isn't showing up in the output of make. This may indicate the wrong shell - type. The .bashrc file sets the environment for the Bourne shell, but - the above error could indicate that the c- shell is being used instead. + It isn't making its way into the configuration because BUFRLIB_v11.3.0 + isn't showing up in the output of make. This may indicate the wrong shell + type. The .bashrc file sets the environment for the Bourne shell, but + the above error could indicate that the C shell is being used instead. - Try the following 2 things: + Try the following 2 things: - 1. Check to make sure this file exists: + 1. Check to make sure this file exists: - .. code-block:: none + .. code-block:: none - ls /home/username/BUFRLIB_v11.3.0/libbufr.a + ls /home/username/BUFRLIB_v11.3.0/libbufr.a - 2. Rerun the MET configure command using the following option on the - command line: + 2. Rerun the MET configure command using the following option on the + command line: - .. code-block:: none + .. code-block:: none - MET_BUFRLIB=/home/username/BUFRLIB_v11.3.0 + MET_BUFRLIB=/home/username/BUFRLIB_v11.3.0 - After doing that, please try recompiling MET. If it fails, please - submit the following log files: "make_install.log" as well as - "config.log" with a new post in the - `METplus GitHub Discussions Forum `_. + After doing that, please try recompiling MET. If it fails, please + submit the following log files: "make_install.log" as well as + "config.log" with a new post in the + `METplus GitHub Discussions Forum `_. Command Line Double Quotes -------------------------- - .. dropdown:: Troubleshooting Help +.. dropdown:: Troubleshooting Help - Single quotes, double quotes, and escape characters can be difficult for - MET to parse. If there are problems, especially in Python code, try - breaking the command up like the below example. + Single quotes, double quotes, and escape characters can be difficult for + MET to parse. If there are problems, especially in Python code, try + breaking the command up like the below example. - .. code-block:: none + .. code-block:: none - ['regrid_data_plane', - '/h/data/global/WXQC/data/umm/1701150006', - 'G003', '/h/data/global/WXQC/data/met/nc_mdl/umm/1701150006', '- field', - '\'name="HGT"; level="P500";\'', '-v', '6'] + ['regrid_data_plane', + '/h/data/global/WXQC/data/umm/1701150006', + 'G003', '/h/data/global/WXQC/data/met/nc_mdl/umm/1701150006', '-field', + '\'name="HGT"; level="P500";\'', '-v', '6'] Environment Variable Settings ----------------------------- - .. dropdown:: Troubleshooting Help +.. dropdown:: Troubleshooting Help - In the below incorrect example for many environment variables have both - the main variable set and the INC and LIB variables set: + The incorrect example below, which applies to many environment variables, + has both the main variable and the INC and LIB variables set: - .. code-block:: none + .. code-block:: none - export MET_GSL=$MET_LIB_DIR/gsl - export MET_GSLINC=$MET_LIB_DIR/gsl/include/gsl - export MET_GSLLIB=$MET_LIB_DIR/gsl/lib + export MET_GSL=$MET_LIB_DIR/gsl + export MET_GSLINC=$MET_LIB_DIR/gsl/include/gsl + export MET_GSLLIB=$MET_LIB_DIR/gsl/lib - **only MET_GSL *OR *MET_GSLINC *AND *MET_GSLLIB need to be set.** - So, for example, either set: + **only MET_GSL OR MET_GSLINC AND MET_GSLLIB need to be set.** + So, for example, either set: - .. code-block:: none + .. code-block:: none - export MET_GSL=$MET_LIB_DIR/gsl + export MET_GSL=$MET_LIB_DIR/gsl - or set: + or set: - .. code-block:: none + .. code-block:: none - export MET_GSLINC=$MET_LIB_DIR/gsl/include/gsl export MET_GSLLIB=$MET_LIB_DIR/gsl/lib + export MET_GSLINC=$MET_LIB_DIR/gsl/include/gsl export MET_GSLLIB=$MET_LIB_DIR/gsl/lib - Additionally, MET does not use MET_HDF5INC and MET_HDF5LIB. - It only uses MET_HDF5. + Additionally, MET does not use MET_HDF5INC and MET_HDF5LIB. + It only uses MET_HDF5. - Our online tutorial can help figure out what should be set and what the - value should be: - https://met.readthedocs.io/en/latest/Users_Guide/installation.html + Our online tutorial can help figure out what should be set and what the + value should be: + https://met.readthedocs.io/en/latest/Users_Guide/installation.html NetCDF Install Issues --------------------- - .. dropdown:: Troubleshooting Help +.. dropdown:: Troubleshooting Help - This example shows a problem with NetCDF in the make_install.log file: + This example shows a problem with NetCDF in the make_install.log file: - .. code-block:: none + .. code-block:: none - /usr/bin/ld: warning: libnetcdf.so.11, - needed by /home/zzheng25/metinstall/lib/libnetcdf_c++4.so, - may conflict with libnetcdf.so.7 + /usr/bin/ld: warning: libnetcdf.so.11, + needed by /home/zzheng25/metinstall/lib/libnetcdf_c++4.so, + may conflict with libnetcdf.so.7 - Below are examples of too many MET_NETCDF options: + Below are examples of too many MET_NETCDF options: - .. code-block:: none + .. code-block:: none - MET_NETCDF='/home/username/metinstall/' - MET_NETCDFINC='/home/username/local/include' - MET_NETCDFLIB='/home/username/local/lib' + MET_NETCDF='/home/username/metinstall/' + MET_NETCDFINC='/home/username/local/include' + MET_NETCDFLIB='/home/username/local/lib' - Either MET_NETCDF **OR** MET_NETCDFINC **AND** MET_NETCDFLIB - need to be set. - If the NetCDF include files are in */home/username/local/include* and the - NetCDF library files are in */home/username/local/lib*, unset the - MET_NETCDF environment variable, then run "make clean", reconfigure, - and then run "make install" and "make test" again. + Either MET_NETCDF **OR** MET_NETCDFINC **AND** MET_NETCDFLIB + need to be set. + If the NetCDF include files are in */home/username/local/include* and the + NetCDF library files are in */home/username/local/lib*, unset the + MET_NETCDF environment variable, then run "make clean", reconfigure, + and then run "make install" and "make test" again. Error While Loading Shared Libraries ------------------------------------ - .. dropdown:: Troubleshooting Help +.. dropdown:: Troubleshooting Help - * Add the lib dir to your LD_LIBRARY_PATH. For example, if you receive - the following error: "./mode_analysis: error while loading shared - libraries: libgsl.so.19: cannot open shared object file: - No such file or directory", you should add the path to the - gsl lib (for example, */home/user/MET/gsl-2.1/lib*) - to your LD_LIBRARY_PATH. + * Add the lib dir to your LD_LIBRARY_PATH. For example, if you receive + the following error: "./mode_analysis: error while loading shared + libraries: libgsl.so.19: cannot open shared object file: + No such file or directory", you should add the path to the + gsl lib (for example, */home/user/MET/gsl-2.1/lib*) + to your LD_LIBRARY_PATH. General Troubleshooting ----------------------- - .. dropdown:: Troubleshooting Help +.. dropdown:: Troubleshooting Help - * For configuration files used, make certain to use empty square brackets - (e.g. [ ]) to indicate no stratification is desired. Do NOT use empty - double quotation marks inside square brackets (e.g. [""]). + * For configuration files used, make certain to use empty square brackets + (e.g., [ ]) to indicate no stratification is desired. Do NOT use empty + double quotation marks inside square brackets (e.g., [""]). - * Have you designated all the required command line arguments? + * Have you designated all the required command line arguments? - * Try rerunning with a higher verbosity level. Increasing the verbosity - level to 4 or 5 prints much more diagnostic information to the screen. + * Try rerunning with a higher verbosity level. Increasing the verbosity + level to 4 or 5 prints much more diagnostic information to the screen. Where to Get Help ================= diff --git a/docs/Users_Guide/appendixB.rst b/docs/Users_Guide/appendixB.rst index ec76da6f23..b9216d7aa0 100644 --- a/docs/Users_Guide/appendixB.rst +++ b/docs/Users_Guide/appendixB.rst @@ -34,7 +34,7 @@ The following map projections are currently supported in MET: Grid Specification Strings ========================== -Several configuration file and command line options support the definition of grids as a grid specification string. A description of the that string for each of the supported grid types is provided below. +Several configuration file and command line options support the definition of grids as a grid specification string. A description of that string for each of the supported grid types is provided below. Lambert Conformal Grid ---------------------- @@ -47,7 +47,7 @@ To specify a Lambert Conformal Grid, the syntax is Here, **Nx** and **Ny** are the number of points in the **x** and **y** grid directions, respectively. These two numbers give the overall size of the grid. **lat_ll** and **lon_ll** are the latitude and longitude, in degrees, of the lower left point of the grid. North latitude and east longitude are considered positive. **lon_orient** is the orientation longitude of the grid. It's the meridian of longitude that's parallel to one of the vertical grid directions. **D_km** and **R_km** are the grid resolution and the radius of the Earth, both in kilometers. **standard_lat_1** and **standard_lat_2** are the standard parallels of the Lambert projection. If the two latitudes are the same, then only one needs to be given. **N|S** means to write either **N** or **S** depending on whether the Lambert projection is from the north pole or the south pole. -As an example of specifying a Lambert grid, suppose you have a northern hemisphere Lambert grid with 614 points in the x direction and 428 points in the y direction. The lower left corner of the grid is at latitude :math:`12.190^\circ` north and longitude :math:`133.459^\circ` west. The orientation longitude is :math:`95^\circ` west. The grid spacing is :math:`12.19058^\circ` km. The radius of the Earth is the default value used in many grib files: 6367.47 km. Both standard parallels are at :math:`25^\circ` north. To specify this grid in the config file, you would write +As an example of specifying a Lambert grid, suppose you have a northern hemisphere Lambert grid with 614 points in the x direction and 428 points in the y direction. The lower left corner of the grid is at latitude :math:`12.190^\circ` north and longitude :math:`133.459^\circ` west. The orientation longitude is :math:`95^\circ` west. The grid spacing is 12.19058 km. The radius of the Earth is the default value used in many GRIB files: 6367.47 km. Both standard parallels are at :math:`25^\circ` north. To specify this grid in the config file, you would write .. code-block:: none @@ -62,7 +62,7 @@ To specify a Lambert Azimuthal Equal Area grid, the syntax is laea Nx Ny lat_first lon_first central_lon Dx_km Dy_km standard_lat equatorial_radius_km [ polar_radius_km ] -Here, **Nx** and **Ny** are the number of points in the **x** and **y** grid directions, respectively. **lat_first** and **lon_first** are the latitude and longitude, in degrees, of the lower left point of the grid. **central_lon** is the orientation longitude of the grid. **Dx_km** and **Dy_km** are the grid resolution in the **x** and **y** directions, both in kilometers. **standard_lat** is the stardard parallel of the Lambert projection. **equatorial_radius_km** is the radius of the Earth at the equator in kilometers. For an elliptical earth, **polar_radius_km** is the radius of the Earth at the poles in kilometers. If both are provided, an elliptical Earth is assumed. If only **equatorial_radius_km** is provided, a spherical Earth is assumed. +Here, **Nx** and **Ny** are the number of points in the **x** and **y** grid directions, respectively. **lat_first** and **lon_first** are the latitude and longitude, in degrees, of the lower left point of the grid. **central_lon** is the orientation longitude of the grid. **Dx_km** and **Dy_km** are the grid resolution in the **x** and **y** directions, both in kilometers. **standard_lat** is the standard parallel of the Lambert projection. **equatorial_radius_km** is the radius of the Earth at the equator in kilometers. For an elliptical earth, **polar_radius_km** is the radius of the Earth at the poles in kilometers. If both are provided, an elliptical Earth is assumed. If only **equatorial_radius_km** is provided, a spherical Earth is assumed. Polar Stereographic Grid @@ -79,7 +79,7 @@ Here, **Nx, Ny, lat_ll, lon_ll, lon_orient, D_km** and **R_km** have the same me Lat/Lon Grid ------------ -For Plate Carrée (i.e. Lat/Lon) grids, the syntax is +For Plate Carrée (i.e., Lat/Lon) grids, the syntax is .. code-block:: none @@ -90,13 +90,13 @@ The parameters **Nx, Ny, lat_ll** and **lon_ll** are as before. **delta_lat** an Rotated Lat/Lon Grid -------------------- -For a Rotated Plate Carrée (i.e. Rotated Lat/Lon) grids, the syntax is +For a Rotated Plate Carrée (i.e., Rotated Lat/Lon) grid, the syntax is .. code-block:: none rotlatlon Nx Ny lat_ll lon_ll delta_lat delta_lon true_lat_sp true_lon_sp aux_rotation -The parameters **Nx, Ny, lat_ll, lon_ll, delta_lat,** and **delta_lon** are as before. **true_lat_sp** and **true_lon_sp** are the latitude and longitude for the south pole. **aux_rotation** is the auxilary rotation in degrees. +The parameters **Nx, Ny, lat_ll, lon_ll, delta_lat,** and **delta_lon** are as before. **true_lat_sp** and **true_lon_sp** are the latitude and longitude for the south pole. **aux_rotation** is the auxiliary rotation in degrees. Mercator Grid ------------- @@ -131,12 +131,12 @@ For a Range/Azimuth grid, the syntax is rngazi range_n azimuth_n range_max_km lat_center lon_center -The parameters **lat_center** and **lon_center** define the latitude and longitude for the center of the grid. The **range_n** and **max_range_km** parameters define the number of and maximum value of the ranges, with spacing in kilometers defined as **max_range_km** / ( **range_n** - 1 ). The **azimuth_n** parameter defines the number of equally-spaced azimuth values, with spacing of 360 / **azimuth_n** degrees clockwise from due east. +The parameters **lat_center** and **lon_center** define the latitude and longitude for the center of the grid. The **range_n** and **range_max_km** parameters define the number of and maximum value of the ranges, with spacing in kilometers defined as **range_max_km** / ( **range_n** - 1 ). The **azimuth_n** parameter defines the number of equally-spaced azimuth values, with spacing of 360 / **azimuth_n** degrees clockwise from due east. Semi Lat/Lon Grid ----------------- -For a Semi Lat/Lon grid, no grid specification string is supported. This grid type is only supported via Python embedding or when reading NetCDF files generated by another MET tool. A Semi Lat/Lon grid defines the information about 2D field of data whose dimension are defined by arrays of latitude (**lats**), longitude (**lons**), level (**levels**), and time (**times**). Times are defined as unixtime, the number of seconds since January 1, 1970. Typically, the lats or lons array and the levels or times array has non-zero length. For example, a zonal mean field is defined using the lats and levels array. A meridional mean field is defined using the lons and levels array. A Hovmoeller field is defined using lats or lons versus times. An arbitrary cross-section is defined by specifying both the lats and lons array with exactly the same length versus levels or times. +For a Semi Lat/Lon grid, no grid specification string is supported. This grid type is only supported via Python embedding or when reading NetCDF files generated by another MET tool. A Semi Lat/Lon grid defines the information about 2D field of data whose dimensions are defined by arrays of latitude (**lats**), longitude (**lons**), level (**levels**), and time (**times**). Times are defined as unixtime, the number of seconds since January 1, 1970. Typically, the lats or lons array and the levels or times array has non-zero length. For example, a zonal mean field is defined using the lats and levels array. A meridional mean field is defined using the lons and levels array. A Hovmoeller field is defined using lats or lons versus times. An arbitrary cross-section is defined by specifying both the lats and lons array with exactly the same length versus levels or times. Statistics can be computed from data on Semi Lat/Lon grids but only when all data resides on the same Semi Lat/Lon grid. Two Semi Lat/Lon grids are equal when their lats, lons, levels, and times arrays match. No functionality is provided to regrid Semi Lat/Lon data. The MET tools can plot Semi Lat/Lon data, however no map data is overlaid since these grids lack two spatial dimensions. diff --git a/docs/Users_Guide/appendixC.rst b/docs/Users_Guide/appendixC.rst index de4c59d961..1d8e4f2c52 100644 --- a/docs/Users_Guide/appendixC.rst +++ b/docs/Users_Guide/appendixC.rst @@ -25,7 +25,7 @@ Which statistics are the same, but with different names? * - Gilbert Skill Score - Equitable Threat Score * - Hanssen and Kuipers Discriminant - - True Skill Statistic, Pierce's Skill Score + - True Skill Statistic, Peirce's Skill Score * - Heidke Skill Score - Cohen's K * - Odds Ratio Skill Score @@ -272,7 +272,7 @@ Instead of C2 being calculated by the user’s dataset, .. math:: \text{HSS } = \text{T*EC }, where EC is allowed to be prescribed by the user ranging from 0 to 1. By default the EC is set to 1 divided by the number of contingency table categories, -e.g. EC is set to 0.33333 for a 3 category (tercile) forecast and 0.5 for a two category (binary) forecast. +e.g., EC is set to 0.33333 for a 3 category (tercile) forecast and 0.5 for a two category (binary) forecast. HSS_EC can range from minus infinity to 1. A perfect forecast would have HSS_EC = 1. @@ -327,7 +327,7 @@ The extreme dependency index measures the association between forecast and obser where *H* and *F* are the Hit Rate and False Alarm Rate, respectively. -EDI can range from :math:`-\infty` to 1, with 0 representing no skill. A perfect forecast would have a value of EDI = 1 (:ref:`Ferro and Stephenson, 2011 `). +EDI can range from :math:`-\infty` to 1, with 0 representing no skill. A perfect forecast would have a value of EDI = 1 (:ref:`Ferro and Stephenson, 2011 `). Symmetric Extreme Dependency Score (SEDS) ----------------------------------------- @@ -338,7 +338,7 @@ The symmetric extreme dependency score measures the association between forecast .. math:: \text{SEDS } = \frac{2 \ln [\frac{(n_{11} + n_{01}) (n_{11} + n_{10})}{T^2}]}{\ln (\frac{n_{11}}{T})} - 1. -SEDS can range from :math:`-\infty` to 1, with 0 representing no skill. A perfect forecast would have a value of SEDS = 1 (:ref:`Ferro and Stephenson, 2011 `). +SEDS can range from :math:`-\infty` to 1, with 0 representing no skill. A perfect forecast would have a value of SEDS = 1 (:ref:`Ferro and Stephenson, 2011 `). Symmetric Extremal Dependency Index (SEDI) ------------------------------------------ @@ -351,7 +351,7 @@ The symmetric extremal dependency index measures the association between forecas where :math:`H = \frac{n_{11}}{n_{11} + n_{01}}` and :math:`F = \frac{n_{10}}{n_{00} + n_{10}}` are the Hit Rate and False Alarm Rate, respectively. -SEDI can range from :math:`-\infty` to 1, with 0 representing no skill. A perfect forecast would have a value of SEDI = 1. SEDI approaches 1 only as the forecast approaches perfection (:ref:`Ferro and Stephenson, 2011 `). +SEDI can range from :math:`-\infty` to 1, with 0 representing no skill. A perfect forecast would have a value of SEDI = 1. SEDI approaches 1 only as the forecast approaches perfection (:ref:`Ferro and Stephenson, 2011 `). Bias-Adjusted Gilbert Skill Score (BAGSS) ----------------------------------------- @@ -383,11 +383,11 @@ Included in SEEPS output :numref:`table_PS_format_info_SEEPS` and SEEPS_MPR outp The SEEPS scoring matrix (equation 15 from :ref:`Rodwell et al, 2010 `) is: .. math:: \{S^{S}_{vf}\} = \frac{1}{2} - \begin{Bmatrix} - 0 & \frac{1}{1-p_1} & \frac{1}{p_3} + \frac{1}{1-p_1}\\ - \frac{1}{p_1} & 0 & \frac{1}{p_3}\\ - \frac{1}{p_1} + \frac{1}{1-p_3} & \frac{1}{1-p_3} & 0 - \end{Bmatrix} + \begin{Bmatrix} + 0 & \frac{1}{1-p_1} & \frac{1}{p_3} + \frac{1}{1-p_1}\\ + \frac{1}{p_1} & 0 & \frac{1}{p_3}\\ + \frac{1}{p_1} + \frac{1}{1-p_3} & \frac{1}{1-p_3} & 0 + \end{Bmatrix} In addition, Rodwell et al (2011) note that SEEPS can be written as the mean of two 2-category scores that individually assess the dry/light and light/heavy thresholds (:ref:`Rodwell et al., 2011 `). Each of these scores is like 1 – HK, but written as: @@ -463,7 +463,7 @@ Called "SP_CORR" in CNT :numref:`table_PS_format_info_CNT` The Spearman rank correlation coefficient (:math:`\rho_{s}`) is a robust measure of association that is based on the ranks of the forecast and observed values rather than the actual values. That is, the forecast and observed samples are ordered from smallest to largest and rank values (from 1 to **n**, where **n** is the total number of pairs) are assigned. The pairs of forecast-observed ranks are then used to compute a correlation coefficient, analogous to the Pearson correlation coefficient, **r**. -A simpler formulation of the Spearman-rank correlation is based on differences between the each of the pairs of ranks (denoted as :math:`d_{i}`): +A simpler formulation of the Spearman-rank correlation is based on differences between each of the pairs of ranks (denoted as :math:`d_{i}`): .. math:: \rho_{s} = \frac{6}{n(n^2 - 1)} \sum_{i=1}^n d_i^2 @@ -478,7 +478,7 @@ Kendall's Tau statistic ( :math:`\tau`) is a robust measure of the level of asso .. math:: \tau = \frac{N_C - N_D}{n(n - 1) / 2} -where :math:`N_C` is the number of "concordant" pairs and :math:`N_D` is the number of "discordant" pairs. Concordant pairs are identified by comparing each pair with all other pairs in the sample; this can be done most easily by ordering all of the ( :math:`f_{i}, o_{i}`) pairs according to :math:`f_{i}`, in which case the :math:`o_{i}` values won't necessarily be in order. The number of concordant matches of a particular pair with other pairs is computed by counting the number of pairs (with larger values) for which the value of :math:`o_i` for the current pair is exceeded (that is, pairs for which the values of **f** and **o** are both larger than the value for the current pair). Once this is done, :math:`N_C` is computed by summing the counts for all pairs. The total number of possible pairs is :math:`N_C`; thus, the number of discordant pairs is :math:`N_D`. +where :math:`N_C` is the number of "concordant" pairs and :math:`N_D` is the number of "discordant" pairs. Concordant pairs are identified by comparing each pair with all other pairs in the sample; this can be done most easily by ordering all of the ( :math:`f_{i}, o_{i}`) pairs according to :math:`f_{i}`, in which case the :math:`o_{i}` values won't necessarily be in order. The number of concordant matches of a particular pair with other pairs is computed by counting the number of pairs (with larger values) for which the value of :math:`o_i` for the current pair is exceeded (that is, pairs for which the values of **f** and **o** are both larger than the value for the current pair). Once this is done, :math:`N_C` is computed by summing the counts for all pairs. The total number of possible pairs is :math:`n(n - 1) / 2`; thus, the number of discordant pairs is :math:`N_D = n(n - 1) / 2 - N_C`. Like **r** and :math:`\rho_{s}`, Kendall's Tau ( :math:`\tau`) ranges between -1 and 1; a value of 1 indicates perfect association (concordance) and a value of -1 indicates perfect negative association. A value of 0 indicates that the forecasts and observations are not associated. @@ -566,7 +566,7 @@ MAE is less influenced by large errors and also does not depend on the mean erro InterQuartile Range of the Errors (IQR) --------------------------------------- -Called "IQR" in CNT output :numref:`table_PS_format_info_CNT` +Called "EIQR" in CNT output :numref:`table_PS_format_info_CNT` The InterQuartile Range of the Errors (IQR) is the difference between the 75th and 25th percentiles of the errors. It is defined as :math:`\text{IQR} = p_{75} (f_i - o_i) - p_{25} (f_i - o_i)`. @@ -641,7 +641,7 @@ Partial Sums Lines (SL1L2, SAL1L2, VL1L2, VAL1L2) :numref:`table_PS_format_info_SL1L2`, :numref:`table_PS_format_info_SAL1L2`, :numref:`table_PS_format_info_VL1L2`, and :numref:`table_PS_format_info_VAL1L2` -The SL1L2, SAL1L2, VL1L2, and VAL1L2 line types are used to store data summaries (e.g. partial sums) that can later be accumulated into verification statistics. These are divided according to scalar or vector summaries (S or V). The climate anomaly values (A) can be stored in place of the actuals, which is just a re-centering of the values around the climatological average. L1 and L2 refer to the L1 and L2 norms, the distance metrics commonly referred to as the "city block" and "Euclidean" distances. The city block is the absolute value of a distance while the Euclidean distance is the square root of the squared distance. +The SL1L2, SAL1L2, VL1L2, and VAL1L2 line types are used to store data summaries (e.g., partial sums) that can later be accumulated into verification statistics. These are divided according to scalar or vector summaries (S or V). The climate anomaly values (A) can be stored in place of the actuals, which is just a re-centering of the values around the climatological average. L1 and L2 refer to the L1 and L2 norms, the distance metrics commonly referred to as the "city block" and "Euclidean" distances. The city block is the absolute value of a distance while the Euclidean distance is the square root of the squared distance. The partial sums can be accumulated over individual cases to produce statistics for a longer period without any loss of information because these sums are *sufficient* for resulting statistics such as RMSE, bias, correlation coefficient, and MAE (:ref:`Mood et al., 1974 `). Thus, the individual errors need not be stored, all of the information relevant to calculation of statistics are contained in the sums. As an example, the sum of all data points and the sum of all squared data points (or equivalently, the sample mean and sample variance) are *jointly sufficient* for estimates of the Gaussian distribution mean and variance. @@ -741,7 +741,7 @@ Gradient Values Called "TOTAL", "FGBAR", "OGBAR", "MGBAR", "EGBAR", "S1", "S1_OG", "FGOG_RATIO", "FGMAG", "OGMAG", "MAG_RMSE", and "LAPLACE_RMSE" in GRAD output :numref:`table_GS_format_info_GRAD` -These statistics are only computed by the Grid-Stat tool and require vectors. Here :math:`\nabla` is the gradient operator, which in this applications signifies the difference between adjacent grid points in both the grid-x and grid-y directions. TOTAL is the count of grid locations used in the calculations. The remaining measures are defined below: +These statistics are only computed by the Grid-Stat tool and require vectors. Here :math:`\nabla` is the gradient operator, which in this application signifies the difference between adjacent grid points in both the grid-x and grid-y directions. TOTAL is the count of grid locations used in the calculations. The remaining measures are defined below: .. math:: \text{FGBAR} = \text{Mean}|\nabla f| = \frac{1}{n} \sum_{i=1}^n | \nabla f_i| @@ -891,7 +891,7 @@ The Brier score is the mean squared probability error. In MET, the Brier Score ( .. math:: \text{BS} = \frac{1}{T} \sum_{i=1}^K [n_{i1} (1 - p_i)^2 + n_{i0} p_i^2] -The equation you will most often see in references uses the individual probability forecasts ( :math:`\rho_{i}`) and the corresponding observations ( :math:`o_{i}`), and is given as :math:`\text{BS} = \frac{1}{T}\sum (p_i - o_i)^2`. This equation is equivalent when the midpoints of the binned probability values are used as the :math:`p_i` . +The equation you will most often see in references uses the individual probability forecasts ( :math:`p_{i}`) and the corresponding observations ( :math:`o_{i}`), and is given as :math:`\text{BS} = \frac{1}{T}\sum (p_i - o_i)^2`. This equation is equivalent when the midpoints of the binned probability values are used as the :math:`p_i` . BS can be partitioned into three terms: (1) reliability, (2) resolution, and (3) uncertainty (:ref:`Murphy, 1987 `). @@ -933,7 +933,7 @@ Calibration Called "CALIBRATION" in PJC output :numref:`table_PS_format_info_PJC` -Calibration is the conditional probability of an event given each probability forecast category (i.e. each row in the **nx2** contingency table). This set of measures is paired with refinement in the calibration-refinement factorization discussed in :ref:`Wilks, 2011 `. A well-calibrated forecast will have calibration values that are near the forecast probability. For example, a 50% probability of precipitation should ideally have a calibration value of 0.5. If the calibration value is higher, then the probability has been underestimated, and vice versa. +Calibration is the conditional probability of an event given each probability forecast category (i.e., each row in the **nx2** contingency table). This set of measures is paired with refinement in the calibration-refinement factorization discussed in :ref:`Wilks, 2011 `. A well-calibrated forecast will have calibration values that are near the forecast probability. For example, a 50% probability of precipitation should ideally have a calibration value of 0.5. If the calibration value is higher, then the probability has been underestimated, and vice versa. .. math:: \text{Calibration}(i) = \frac{n_{i1}}{n_{1.}} = \text{probability}(o_1|p_i) @@ -962,7 +962,7 @@ Base Rate Called "BASER" in PJC output :numref:`table_PS_format_info_PJC` -This is the probability of an event for each forecast category :math:`p_i` (row), i.e. the conditional base rate. This set of measures is paired with likelihood in the likelihood-base rate factorization, see :ref:`Wilks, 2011 ` for further information. This measure is calculated for each row of the contingency table. Ideally, the event should become more frequent as the probability forecast increases. +This is the probability of an event for each forecast category :math:`p_i` (row), i.e., the conditional base rate. This set of measures is paired with likelihood in the likelihood-base rate factorization, see :ref:`Wilks, 2011 ` for further information. This measure is calculated for each row of the contingency table. Ideally, the event should become more frequent as the probability forecast increases. .. math:: \text{Base Rate}(i) = \frac{n_{i1}}{n_{i.}} = \text{probability}(o_{i1}) @@ -977,7 +977,7 @@ The ideal forecast (i.e., one with perfect reliability) has conditional observed .. figure:: figure/appendixC-rel_diag.jpg - Example of Reliability Diagram + Example of Reliability Diagram Receiver Operating Characteristic --------------------------------- @@ -992,7 +992,7 @@ A ROC curve shows how well the forecast discriminates between two outcomes, so i .. figure:: figure/appendixC-roc_example.jpg - Example of ROC Curve + Example of ROC Curve Area Under the ROC Curve (AUC) ------------------------------ @@ -1013,7 +1013,7 @@ MET Verification Measures for Ensemble Forecasts RPS --- -Called "RPS" in RPS output :numref:`table_ES_header_info_es_out_ECNT` +Called "RPS" in RPS output :numref:`table_ES_header_info_es_out_RPS` While the above probabilistic verification measures utilize dichotomous observations, the Ranked Probability Score (RPS, :ref:`Epstein, 1969 `, :ref:`Murphy, 1969 `) is the only probabilistic verification measure for discrete multiple-category events available in MET. It is assumed that the categories are ordinal as nominal categorical variables can be collapsed into sequences of binary predictands, which can in turn be evaluated with the above measures for dichotomous variables (:ref:`Wilks, 2011 `). The RPS is the multi-category extension of the Brier score (:ref:`Tödter and Ahrens, 2012 `), and is a proper score (:ref:`Mason, 2008 `). @@ -1026,7 +1026,7 @@ To clarify, :math:`F_1 = f_1` is the first component of :math:`F_m`, :math:`F_2 .. math:: \text{RPS} = \sum_{m=1}^J (F_m - O_m)^2 = \sum_{m=1}^J BS_m, -where :math:`BS_m` is the Brier score for the m-th category (:ref:`Tödter and Ahrens, 2012 `). Subsequently, the RPS lends itself to a decomposition into reliability, resolution and uncertainty components, noting that each component is aggregated over the different categories; these are written to the columns named "RPS_REL", "RPS_RES" and "RPS_UNC" in RPS output :numref:`table_ES_header_info_es_out_ECNT`. +where :math:`BS_m` is the Brier score for the m-th category (:ref:`Tödter and Ahrens, 2012 `). Subsequently, the RPS lends itself to a decomposition into reliability, resolution and uncertainty components, noting that each component is aggregated over the different categories; these are written to the columns named "RPS_REL", "RPS_RES" and "RPS_UNC" in RPS output :numref:`table_ES_header_info_es_out_RPS`. CRPS ---- @@ -1041,7 +1041,7 @@ Closed form expressions for the CRPS are difficult to define when using data rat .. math:: \text{crps}_i (N( \mu, \sigma^2),y) = \sigma ( \frac{y - \mu}{\sigma} (2 \Phi (\frac{y - \mu}{\sigma}) -1) + 2 \phi (\frac{y - \mu}{\sigma}) - \frac{1}{\sqrt{\pi}}) -In this equation, the y represents the event threshold. The estimated mean and standard deviation of the ensemble forecasts ( :math:`\mu \text{ and } \sigma`) are used as the parameters of the normal distribution. The values of the normal distribution are represented by the probability density function (PDF) denoted by :math:`\Phi` and the cumulative distribution function (CDF), denoted in the above equation by :math:`\phi`. +In this equation, the y represents the event threshold. The estimated mean and standard deviation of the ensemble forecasts ( :math:`\mu \text{ and } \sigma`) are used as the parameters of the normal distribution. The values of the normal distribution are represented by the probability density function (PDF) denoted by :math:`\phi` and the cumulative distribution function (CDF), denoted in the above equation by :math:`\Phi`. The overall CRPS is calculated as the average of the individual measures. In equation form: @@ -1122,7 +1122,7 @@ The bias ratio (BIAS_RATIO) is computed when verifying an ensemble against gridd .. math:: \text{BIAS_RATIO} = \frac{ \text{ME}_{f >= o} }{ |\text{ME}_{f < o}| } -A perfect forecast has ME = 0. Since BIAS_RATIO is computed as the high bias (ME_GE_OBS) divide by the absolute value of the low bias (ME_LT_OBS), a perfect forecast has BIAS_RATIO = 0/0, which is undefined. In practice, the high and low bias values are unlikely to be 0. +A perfect forecast has ME = 0. Since BIAS_RATIO is computed as the high bias (ME_GE_OBS) divided by the absolute value of the low bias (ME_LT_OBS), a perfect forecast has BIAS_RATIO = 0/0, which is undefined. In practice, the high and low bias values are unlikely to be 0. The range for BIAS_RATIO is 0 to infinity. A score of 1 indicates that the high and low biases are equal. A score greater than 1 indicates that the high bias is larger than the magnitude of the low bias. A score less than 1 indicates the opposite behavior. @@ -1131,7 +1131,7 @@ IGN Called "IGN" in ECNT output :numref:`table_ES_header_info_es_out_ECNT` -The ignorance score (IGN) is the negative logarithm of a predictive probability density function (:ref:`Gneiting et al., 2004 `). In MET, the IGN is calculated based on a normal approximation to the forecast distribution (i.e. a normal pdf is fit to the forecast values). This approximation may not be valid, especially for discontinuous forecasts like precipitation, and also for very skewed forecasts. For a single normal distribution **N** with parameters :math:`\mu \text{ and } \sigma`, the ignorance score is +The ignorance score (IGN) is the negative logarithm of a predictive probability density function (:ref:`Gneiting et al., 2004 `). In MET, the IGN is calculated based on a normal approximation to the forecast distribution (i.e., a normal pdf is fit to the forecast values). This approximation may not be valid, especially for discontinuous forecasts like precipitation, and also for very skewed forecasts. For a single normal distribution **N** with parameters :math:`\mu \text{ and } \sigma`, the ignorance score is .. math:: \text{ign} (N( \mu, \sigma),y) = \frac{1}{2} \ln (2 \pi \sigma^2 ) + \frac{(y - \mu)^2}{2\sigma^2}. @@ -1149,7 +1149,7 @@ Observation Error Logarithmic Scoring Rules Called "IGN_CONV_OERR" and "IGN_CORR_OERR" in ECNT output :numref:`table_ES_header_info_es_out_ECNT` -One approach that is used to take observation error into account in a summary measure is to add error to the forecast by a convolution with the observation model (e.g., :ref:`Anderson, 1996 `; :ref:`Hamill, 2001 `; :ref:`Saetra et. al., 2004 `; :ref:`Bröcker and Smith, 2007 `; :ref:`Candille et al., 2007 `; :ref:`Candille and Talagrand, 2008 `; :ref:`Röpnack et al., 2013 `). Specifically, suppose :math:`y=x+w`, where :math:`y` is the observed value, :math:`x` is the true value, and :math:`w` is the error. Then, if :math:`f` is the density forecast for :math:`x` and :math:`\nu` is the observation model, then the implied density forecast for :math:`y` is given by the convolution: +One approach that is used to take observation error into account in a summary measure is to add error to the forecast by a convolution with the observation model (e.g., :ref:`Anderson, 1996 `; :ref:`Hamill, 2001 `; :ref:`Saetra et al., 2004 `; :ref:`Bröcker and Smith, 2007 `; :ref:`Candille et al., 2007 `; :ref:`Candille and Talagrand, 2008 `; :ref:`Röpnack et al., 2013 `). Specifically, suppose :math:`y=x+w`, where :math:`y` is the observed value, :math:`x` is the true value, and :math:`w` is the error. Then, if :math:`f` is the density forecast for :math:`x` and :math:`\nu` is the observation model, then the implied density forecast for :math:`y` is given by the convolution: .. math:: (f*\nu)(y) = \int\nu(y|x)f(x)dx @@ -1163,7 +1163,7 @@ One approach that is used to take observation error into account in a summary me .. math:: \text{IGN_CONV_OERR} = s(f,y) = \frac{1}{2}\log(2 \pi (\sigma^2 + c^2)) + \frac{(y - \mu)^2}{2 (\sigma^2 + c^2)} -Another approach to incorporation of observation uncertainty into a measure is the error-correction approach. The approach merely ensures that the scoring rule, :math:`s`, is unbiased for a scoring rule :math:`s_0` if they have the same expected value. :ref:`Ferro, 2017 ` gives the error-corrected ignorance scoring rule (which is also proposer when :math:`w\sim N(0,c^2)`) as +Another approach to incorporation of observation uncertainty into a measure is the error-correction approach. The approach merely ensures that the scoring rule, :math:`s`, is unbiased for a scoring rule :math:`s_0` if they have the same expected value. :ref:`Ferro, 2017 ` gives the error-corrected ignorance scoring rule (which is also proper when :math:`w\sim N(0,c^2)`) as .. only:: latex @@ -1262,19 +1262,19 @@ Uniform Fractions Skill Score Called "UFSS" in NBRCNT output :numref:`table_GS_format_info_NBRCNT` -The Uniform Fractions Skill Score (UFSS) is a reference statistic for the Fractions Skill score based on a uniform distribution of the total observed events across the grid. UFSS represents the FSS that would be obtained at the grid scale from a forecast with a fraction/probability equal to the total observed event proportion at every point. The formula is :math:`UFSS = (1 + f_o)/2` (i.e., halfway between perfect skill and random forecast skill) where :math:`f_o` is the total observed event proportion (i.e. observation rate). +The Uniform Fractions Skill Score (UFSS) is a reference statistic for the Fractions Skill score based on a uniform distribution of the total observed events across the grid. UFSS represents the FSS that would be obtained at the grid scale from a forecast with a fraction/probability equal to the total observed event proportion at every point. The formula is :math:`UFSS = (1 + f_o)/2` (i.e., halfway between perfect skill and random forecast skill) where :math:`f_o` is the total observed event proportion (i.e., observation rate). Forecast Rate ------------- -Called "F_rate" in NBRCNT output :numref:`table_GS_format_info_NBRCNT` +Called "F_RATE" in NBRCNT output :numref:`table_GS_format_info_NBRCNT` The overall proportion of grid points with forecast events to total grid points in the domain. The forecast rate will match the observation rate in unbiased forecasts. Observation Rate ---------------- -Called "O_rate" in NBRCNT output :numref:`table_GS_format_info_NBRCNT` +Called "O_RATE" in NBRCNT output :numref:`table_GS_format_info_NBRCNT` The overall proportion of grid points with observed events to total grid points in the domain. The forecast rate will match the observation rate in unbiased forecasts. This quantity is sometimes referred to as the base rate. @@ -1330,9 +1330,9 @@ Unlike Baddeley's :math:`\Delta` metric, the MED is not a mathematical metric be .. math:: min \text{MED}(A,B) = min( \text{MED}(A,B),\text{MED}(B,A)) - max \text{MED}(A,B) = max( \text{MED}(A,B), \text{MED}(B,A)) + max \text{MED}(A,B) = max( \text{MED}(A,B), \text{MED}(B,A)) - mean \text{MED}(A,B) = \frac{1}{2}(\text{MED}(A,B) + \text{MED}(B,A)) + mean \text{MED}(A,B) = \frac{1}{2}(\text{MED}(A,B) + \text{MED}(B,A)) From the distance map perspective, MED *(A,B)* is the average of the values in :numref:`grid-stat_fig4` (top right), and MED *(B,A)* is the average of the values in :numref:`grid-stat_fig4` (bottom left). Note that the average is only over the circular regions depicted in the figure. @@ -1400,7 +1400,7 @@ Suppose now that we have a collection of *N* data points :math:`x_i \text{for } .. math:: I = \lfloor (N - 1)t \rfloor - \Delta = (N - 1)t - I + \Delta = (N - 1)t - I Then the value *p* of the percentile is diff --git a/docs/Users_Guide/appendixD.rst b/docs/Users_Guide/appendixD.rst index c398fd3313..a5cd713095 100644 --- a/docs/Users_Guide/appendixD.rst +++ b/docs/Users_Guide/appendixD.rst @@ -16,13 +16,13 @@ The most commonly used confidence interval about an estimate for a statistic (or where :math:`z_{\alpha / 2}` is the :math:`\alpha - \text{th}` quantile of the standard normal distribution, and :math:`V(\theta )` is the standard error of the statistic (or parameter), :math:`\theta`. For example, the most common example is for the mean of a sample, :math:`X_1,\cdots,X_n`, of independent and identically distributed (iid) normal random variables with mean :math:`\mu` and variance :math:`\sigma`. Here, the mean is estimated by :math:`\frac{1}{n} \sum_{i=1}^n X_i = \bar{X}`, and the standard error is just the standard deviation of the random variables divided by the square root of the sample size. That is, :math:`V( \theta ) = V ( \bar{X} ) = \frac{\sigma}{\sqrt{n}}`, and this must be estimated by :math:`\hat{V} (\bar{X} )`, which is obtained here by replacing :math:`\sigma` by its estimate, :math:`\hat{\sigma}`, where :math:`\hat{\sigma} = \frac{1}{n - 1} \sum_{i=1}^n (X_i - \bar{X})^2`. -Mostly, the normal approximation is used as an asymptotic approximation. That is, the interval for :math:`\theta` may only be appropriate for large **n**. For small **n**, the mean has an interval based on the Student's **t** distribution with **n-1** degrees of freedom. Essentially, :math:`z_{\alpha / 2}` of the question is replaced with the quantile of this **t** distribution. That is, the interval is given by +Mostly, the normal approximation is used as an asymptotic approximation. That is, the interval for :math:`\theta` may only be appropriate for large **n**. For small **n**, the mean has an interval based on the Student's **t** distribution with **n-1** degrees of freedom. Essentially, :math:`z_{\alpha / 2}` in the equation above is replaced with the quantile of this **t** distribution. That is, the interval is given by .. math:: \mu \pm t_{\alpha / 2,\nu - 1} \cdot \frac{\sigma}{\sqrt{n}} where again, :math:`\sigma` is replaced by its estimate, :math:`\hat{\sigma}`, as described above. -:numref:`verif_stat_approx_CI` summarizes the verification statistics in MET that have normal approximation CIs given by :math:`\theta` along with their corresponding standard error estimates, . It should be noted that for the first two rows of this table (i.e., Forecast/Observation Mean and Mean error) MET also calculates the interval around :math:`\mu` for small sample sizes. +:numref:`verif_stat_approx_CI` summarizes the verification statistics in MET that have normal approximation CIs given by :math:`\theta` along with their corresponding standard error estimates, :math:`\hat{V}(\theta)`. It should be noted that for the first two rows of this table (i.e., Forecast/Observation Mean and Mean error) MET also calculates the interval around :math:`\mu` for small sample sizes. .. _verif_stat_approx_CI: @@ -82,5 +82,5 @@ Typically, a simple random sample is taken for step 2, and that is how it is don There are numerous ways to construct CIs from the sample obtained in step 4. MET allows for two of these procedures: the percentile and the BCa. The percentile is the most commonly known method, and the simplest to understand. It is merely the :math:`\alpha / 2` and :math:`1 - \alpha / 2` percentiles from the sample of statistics. Unfortunately, however, it has been shown that this interval is too optimistic in practice (i.e., it doesn't have accurate coverage). One solution is to use the BCa method, which is very accurate, but it is also computationally intensive. This method adjusts for bias and non-constant variance, and yields the percentile interval in the event that the sample is unbiased with constant variance. -If there is dependency in the sample, then it is prudent to account for this dependency in some way. :ref:`Gilleland (2010) ` describes the bootstrap procedure, along with the above-mentioned parametric methods, in more detail specifically for the verification application. If there is dependency in the sample, then it is prudent to account for this dependency in some way (see :ref:`Gilleland (2020, part I) ` part I for an in-depth discussion of bootstrapping in the competing forecast verification domain). One method that is particularly appropriate for serially dependent data is the circular block resampling procedure for step 2. +If there is dependency in the sample, then it is prudent to account for this dependency in some way. :ref:`Gilleland (2010) ` describes the bootstrap procedure, along with the above-mentioned parametric methods, in more detail specifically for the verification application. See :ref:`Gilleland (2020, part I) ` for an in-depth discussion of bootstrapping in the competing forecast verification domain. One method that is particularly appropriate for serially dependent data is the circular block resampling procedure for step 2. diff --git a/docs/Users_Guide/appendixE.rst b/docs/Users_Guide/appendixE.rst index 374fdbc330..8827f8a038 100644 --- a/docs/Users_Guide/appendixE.rst +++ b/docs/Users_Guide/appendixE.rst @@ -4,7 +4,7 @@ Appendix E WWMCA Tools ********************** -There are two WWMCA tools available. The WWMCA-Plot tool makes a PostScript plot of one or more WWMCA cloud percent files and the WWMCA-Regrid tool regrids WWMCA cloud percent files and reformats them into netCDF files that the other MET tools can read. +There are two WWMCA tools available. The WWMCA-Plot tool makes a PostScript plot of one or more WWMCA cloud percent files and the WWMCA-Regrid tool regrids WWMCA cloud percent files and reformats them into NetCDF files that the other MET tools can read. The WWMCA tools get valid time and hemisphere (north or south) information from the file names, so it's important for both of the WWMCA tools that these file names not be changed. @@ -24,15 +24,15 @@ The optional **-outdir** argument specifies a directory where the output PostScr .. figure:: figure/reformat_grid_fig2.png - Example output of WWMCA-Plot tool. + Example output of WWMCA-Plot tool. The usage statement for wwmca_regrid is .. code-block:: none - wwmca_regrid -out filename config filename [ -nh filename ] [ -sh filename ] + wwmca_regrid -out filename -config filename [ -nh filename ] [ -sh filename ] -Here, the **-out** switch tells wwmca_regrid what to name the output netCDF file. The **-config** switch gives the name of the config file that wwmca_regrid should use-like many of the MET tools, wwmca-regrid uses a configuration file to specify user-changeable parameters. The format of this file will be explained below. +Here, the **-out** switch tells wwmca_regrid what to name the output NetCDF file. The **-config** switch gives the name of the config file that wwmca_regrid should use-like many of the MET tools, wwmca_regrid uses a configuration file to specify user-changeable parameters. The format of this file will be explained below. The **-nh** and **-sh** options give names of WWMCA cloud percent files that wwmca_regrid should use as input. Northern hemisphere files are specified with **-nh**, and southern hemisphere files with **-sh**. At least one of these must be given, but in many cases both need not be given. @@ -44,7 +44,7 @@ Now let's talk about the details of the config file. The config file has the sam To grid = "G218"; -and that will work. Failing that, you must give the parameters that specify the grid and it's projection. Please refer the description of the grid specification strings in :ref:`appendixB`. +and that will work. Failing that, you must give the parameters that specify the grid and its projection. Please refer to the description of the grid specification strings in :ref:`appendixB`. Thankfully, the rest of the parameters in the config file are easier to specify. @@ -68,9 +68,9 @@ The next variable, **good_percent**, tells what fraction of the values in the in .. code-block:: none - good percent = 0; + good_percent = 0; -The rest of the config file parameters have to do with how the output netCDF file represents the data. These should be self-explanatory, so I'll just give an example: +The rest of the config file parameters have to do with how the output NetCDF file represents the data. These should be self-explanatory, so I'll just give an example: .. code-block:: none diff --git a/docs/Users_Guide/appendixF.rst b/docs/Users_Guide/appendixF.rst index 62221f84d1..13da5d1039 100644 --- a/docs/Users_Guide/appendixF.rst +++ b/docs/Users_Guide/appendixF.rst @@ -36,7 +36,7 @@ Users should be aware that in some cases, the C-language Python header files and The **NumPy**, **Xarray**, and **Pandas** Python packages are required by the Python scripts included with the MET software that facilitate the passing of data in memory. The **SciPy** and **YAML** Python packages are required by the tropical cyclone diagnostics Python scripts called by the TC-Diag tool. The **netCDF4** package is used for reading and writing temporary files for Python embedding, but only when the **MET_PYTHON_TMP_FORMAT** environment variable is set to `netcdf` at runtime. -In addition to using **\-\-enable-python** with **configure** as mentioned above, the following environment variables must also be set prior to executing **configure**: **MET_PYTHON_BIN_EXE**, **MET_PYTHON_CC**, and **MET_PYTHON_LD**. These may either be set as environment variables or as command line options to **configure**. These environment variables are used when building MET to enable the compiler to find the requisite Python executable, header files, and libraries in the user's local filesystem. Fortunately, Python provides a way to set these variables properly. This frees the user from the necessity of having any expert knowledge of the compiling and linking process. Along with the **Python** executable in the users local Python installation, there should be another executable called **python3-config**, whose output can be used to set these environment variables as follows: +In addition to using **\-\-enable-python** with **configure** as mentioned above, the following environment variables must also be set prior to executing **configure**: **MET_PYTHON_BIN_EXE**, **MET_PYTHON_CC**, and **MET_PYTHON_LD**. These may either be set as environment variables or as command line options to **configure**. These environment variables are used when building MET to enable the compiler to find the requisite Python executable, header files, and libraries in the user's local filesystem. Fortunately, Python provides a way to set these variables properly. This frees the user from the necessity of having any expert knowledge of the compiling and linking process. Along with the **Python** executable in the user's local Python installation, there should be another executable called **python3-config**, whose output can be used to set these environment variables as follows: • Set **MET_PYTHON_BIN_EXE** to the full path of the desired Python executable. @@ -49,15 +49,15 @@ Make sure that these are set as environment variables or that you have included If a user attempts to invoke Python embedding with a version of MET that was not compiled with Python, MET will return an ERROR: .. code-block:: none - :caption: MET Errors Without Python Enabled + :caption: MET Errors Without Python Enabled - ERROR : Met2dDataFileFactory::new_met_2d_data_file() -> Support for Python has not been compiled! - ERROR : To run Python scripts, recompile with the --enable-python option. + ERROR : Met2dDataFileFactory::new_met_2d_data_file() -> Support for Python has not been compiled! + ERROR : To run Python scripts, recompile with the --enable-python option. - - or - + - or - - ERROR : process_point_obs() -> Support for Python has not been compiled! - ERROR : To run Python scripts, recompile with the --enable-python option. + ERROR : process_point_obs() -> Support for Python has not been compiled! + ERROR : To run Python scripts, recompile with the --enable-python option. Controlling Which Python MET Uses When Running ============================================== @@ -67,9 +67,9 @@ When MET is compiled with Python embedding support, MET uses the Python executab If a user's Python script requires packages that are not available in the Python installation used when compiling the MET software, they will encounter a runtime error when using MET. In this instance, the user will need to change the Python MET is using to a different installation with the required packages for their script. It is the responsibility of the user to manage this Python installation, and one popular approach is to use a custom Anaconda (Conda) Python environment. Once the Python installation meeting the user's requirements is available, the user can force MET to use it by setting the **MET_PYTHON_EXE** environment variable to the full path of the Python executable in that installation. For example: .. code-block:: none - :caption: Setting MET_PYTHON_EXE + :caption: Setting MET_PYTHON_EXE - export MET_PYTHON_EXE=/usr/local/python3/bin/python3 + export MET_PYTHON_EXE=/usr/local/python3/bin/python3 Setting this environment variable triggers slightly different processing logic in MET than when MET uses the Python installation that was used when compiling MET. When using the Python installation that was used when compiling MET, Python is called directly and data are passed in memory from Python to the MET tools. When the user sets **MET_PYTHON_EXE**, MET does the following: @@ -98,7 +98,7 @@ Details for each of these data structures are provided below. .. note:: - All sample commands and directories listed below are relative to the top level of the MET source code directory. + All sample commands and directories listed below are relative to the top level of the MET source code directory. .. _pyembed-2d-data: @@ -125,58 +125,58 @@ Attributes for 2D Gridded Dataplanes ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ .. list-table:: 2D Dataplane Attributes - :widths: 5 5 10 5 - :header-rows: 1 - - * - key - - description - - data type/format - - required/optional - * - valid - - valid time - - string (YYYYMMDD_HHMMSS) - - required - * - init - - initialization time - - string (YYYYMMDD_HHMMSS) - - required - * - lead - - forecast lead - - string (HHMMSS) - - required - * - accum - - accumulation interval - - string (HHMMSS) - - required - * - name - - variable name - - string - - required - * - long_name - - variable long name - - string - - required - * - level - - variable level - - string - - required - * - units - - variable units - - string - - required - * - grid - - :ref:`grid information` - - string or dict - - required - * - fill_value - - :ref:`missing data value` - - int or float - - optional + :widths: 5 5 10 5 + :header-rows: 1 + + * - key + - description + - data type/format + - required/optional + * - valid + - valid time + - string (YYYYMMDD_HHMMSS) + - required + * - init + - initialization time + - string (YYYYMMDD_HHMMSS) + - required + * - lead + - forecast lead + - string (HHMMSS) + - required + * - accum + - accumulation interval + - string (HHMMSS) + - required + * - name + - variable name + - string + - required + * - long_name + - variable long name + - string + - required + * - level + - variable level + - string + - required + * - units + - variable units + - string + - required + * - grid + - :ref:`grid information` + - string or dict + - required + * - fill_value + - :ref:`missing data value` + - int or float + - optional .. note:: - Often times Xarray DataArray objects come with their own set of attributes available as a property. To avoid conflict with the required attributes - for MET, it is advised to strip these attributes and rely on the **attrs** dictionary defined in your script. + Often times Xarray DataArray objects come with their own set of attributes available as a property. To avoid conflict with the required attributes + for MET, it is advised to strip these attributes and rely on the **attrs** dictionary defined in your script. .. _pyembed-fillvalue-attrs: @@ -190,36 +190,36 @@ Python embedding for 2D gridded dataplanes provides support for a user-defined m If a user has a 2D dataplane with another value that should be considered a fill value by MET, then the user must use the **fill_value** attribute in the **attrs** dictionary. An example would be if a user had a 2D dataplane with missing data indicated with -99. A user can use the **fill_value** attribute in their **attrs** dictionary which will tell MET to ignore those values: .. code-block:: none - :caption: User Fill Value for 2D Dataplane + :caption: User Fill Value for 2D Dataplane - 'fill_value': -99 + 'fill_value': -99 Alternatively, the user can choose to replace their special values with one of the four supported values instead of setting the **fill_value** attribute. Note that only a single user-defined fill value is supported at this time. .. _pyembed-grid-attrs: -The grid entry in the **attrs** dictionary must contain the grid size and projection information in the same format that is used in the netCDF files written out by the MET tools. The value of this item in the dictionary can either be a string, or another dictionary. Examples of the **grid** entry defined as a string are: +The grid entry in the **attrs** dictionary must contain the grid size and projection information in the same format that is used in the NetCDF files written out by the MET tools. The value of this item in the dictionary can either be a string, or another dictionary. Examples of the **grid** entry defined as a string are: • Using a named grid supported by MET: .. code-block:: none - :caption: Named Grid + :caption: Named Grid - 'grid': 'G212' + 'grid': 'G212' • As a grid specification string, as described in :ref:`appendixB`: .. code-block:: none - :caption: Grid Specification String + :caption: Grid Specification String - 'grid': 'lambert 185 129 12.19 -133.459 -95 40.635 6371.2 25 25 N' + 'grid': 'lambert 185 129 12.19 -133.459 -95 40.635 6371.2 25 25 N' • As the path to an existing gridded data file: .. code-block:: none - :caption: Grid From File + :caption: Grid From File - 'grid': '/path/to/sample_data.grib' + 'grid': '/path/to/sample_data.grib' When specified as a dictionary, the contents of the **grid** entry vary based upon the grid **type**. The required elements for supported grid types are: @@ -301,39 +301,39 @@ Additional information about supported grids can be found in :ref:`appendixB`. Finally, an example **attrs** dictionary is shown below: .. code-block:: none - :caption: Sample Attrs Dictionary - - attrs = { - - 'valid': '20050807_120000', - 'init': '20050807_000000', - 'lead': '120000', - 'accum': '120000', - - 'name': 'Foo', - 'long_name': 'FooBar', - 'level': 'Surface', - 'units': 'None', - - # Define 'grid' as a string or a dictionary - - 'grid': { - 'type': 'Lambert Conformal', - 'hemisphere': 'N', - 'name': 'FooGrid', - 'scale_lat_1': 25.0, - 'scale_lat_2': 25.0, - 'lat_pin': 12.19, - 'lon_pin': -135.459, - 'x_pin': 0.0, - 'y_pin': 0.0, - 'lon_orient': -95.0, - 'd_km': 40.635, - 'r_km': 6371.2, - 'nx': 185, - 'ny': 129, - } - } + :caption: Sample Attrs Dictionary + + attrs = { + + 'valid': '20050807_120000', + 'init': '20050807_000000', + 'lead': '120000', + 'accum': '120000', + + 'name': 'Foo', + 'long_name': 'FooBar', + 'level': 'Surface', + 'units': 'None', + + # Define 'grid' as a string or a dictionary + + 'grid': { + 'type': 'Lambert Conformal', + 'hemisphere': 'N', + 'name': 'FooGrid', + 'scale_lat_1': 25.0, + 'scale_lat_2': 25.0, + 'lat_pin': 12.19, + 'lon_pin': -135.459, + 'x_pin': 0.0, + 'y_pin': 0.0, + 'lon_orient': -95.0, + 'd_km': 40.635, + 'r_km': 6371.2, + 'nx': 185, + 'ny': 129, + } + } Running Python Embedding for 2D Gridded Dataplanes ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ @@ -343,27 +343,27 @@ On the command line for any of the MET tools which will be obtaining its data fr Listed below is an example of running the Plot-Data-Plane tool to call a Python script for data that is included with the MET release tarball. Assuming the MET executables are in your path, this example may be run from the top-level MET source code directory: .. code-block:: none - :caption: plot_data_plane Python Embedding + :caption: plot_data_plane Python Embedding - plot_data_plane PYTHON_NUMPY fcst.ps \ - 'name="scripts/python/examples/read_ascii_numpy.py data/python/fcst.txt FCST";' \ - -title "Python enabled plot_data_plane" + plot_data_plane PYTHON_NUMPY fcst.ps \ + 'name="scripts/python/examples/read_ascii_numpy.py data/python/fcst.txt FCST";' \ + -title "Python enabled plot_data_plane" -The first argument for the Plot-Data-Plane tool is the gridded data file to be read. When calling Python script that has a two-dimensional gridded dataplane stored in a NumPy N-D array object, set this to the constant string **PYTHON_NUMPY**. The second argument is the name of the output PostScript file to be written. The third argument is a string describing the data to be plotted. When calling a Python script, set **name** to the full path of the Python script to be run along with any command line arguments for that script. Lastly, the **-title** option is used to add a title to the plot. Note that any print statements included in the Python script will be printed to the screen. The above example results in the following log messages: +The first argument for the Plot-Data-Plane tool is the gridded data file to be read. When calling a Python script that has a two-dimensional gridded dataplane stored in a NumPy N-D array object, set this to the constant string **PYTHON_NUMPY**. The second argument is the name of the output PostScript file to be written. The third argument is a string describing the data to be plotted. When calling a Python script, set **name** to the full path of the Python script to be run along with any command line arguments for that script. Lastly, the **-title** option is used to add a title to the plot. Note that any print statements included in the Python script will be printed to the screen. The above example results in the following log messages: .. code-block:: none - DEBUG 1: Opening data file: PYTHON_NUMPY - Input File: 'data/python/fcst.txt' - Data Name : 'FCST' - Data Shape: (129, 185) - Data Type: dtype('float64') - Attributes: {'name': 'FCST', 'long_name': 'FCST_word', - 'level': 'Surface', 'units': 'None', - 'init': '20050807_000000', 'valid': '20050807_120000', - 'lead': '120000', 'accum': '120000' - 'grid': { ... } } - DEBUG 1: Creating postscript file: fcst.ps + DEBUG 1: Opening data file: PYTHON_NUMPY + Input File: 'data/python/fcst.txt' + Data Name : 'FCST' + Data Shape: (129, 185) + Data Type: dtype('float64') + Attributes: {'name': 'FCST', 'long_name': 'FCST_word', + 'level': 'Surface', 'units': 'None', + 'init': '20050807_000000', 'valid': '20050807_120000', + 'lead': '120000', 'accum': '120000' + 'grid': { ... } } + DEBUG 1: Creating postscript file: fcst.ps .. _met-python-input-arg: @@ -373,40 +373,40 @@ Special Case for Gen-Ens-Prod, Ensemble-Stat, Series-Analysis, and MTD The Gen-Ens-Prod, Ensemble-Stat, Series-Analysis, and MTD tools all have the ability to read multiple input files. Because of this feature, a different approach to Python embedding is required. A typical use of these tools is to provide a list of files on the command line. For example: .. code-block:: - :caption: Gen-Ens-Prod Command Line + :caption: Gen-Ens-Prod Command Line - gen_ens_prod ens1.nc ens2.nc ens3.nc ens4.nc -out ens_prod.nc -config GenEnsProd_config + gen_ens_prod -ens ens1.nc ens2.nc ens3.nc ens4.nc -out ens_prod.nc -config GenEnsProd_config In this case, a user is passing 4 ensemble members to Gen-Ens-Prod to be evaluated, and each member is in a separate file. If a user wishes to use Python embedding to process the ensemble input files, then the same exact command is used; however special modifications inside the GenEnsProd_config file are needed. In the config file dictionary, the user must set the **file_type** entry to either **PYTHON_NUMPY** or **PYTHON_XARRAY** to activate the Python embedding for these tools. Then, in the **name** entry of the config file dictionaries for the forecast or observation data, the user must list the **full path** to the Python script to be run. However, in the Python command, replace the name of the input gridded data file to the Python script with the constant string **MET_PYTHON_INPUT_ARG**. When looping over all of the input files, the MET tools will replace that constant **MET_PYTHON_INPUT_ARG** with the path to the input file currently being processed and optionally, any command line arguments for the Python script. Here is what this looks like in the GenEnsProd_config file for the above example: .. code-block:: - :caption: Gen-Ens-Prod MET_PYTHON_INPUT_ARG Config + :caption: Gen-Ens-Prod MET_PYTHON_INPUT_ARG Config - file_type = PYTHON_NUMPY; - field = [ { name = "gen_ens_prod_pyembed.py MET_PYTHON_INPUT_ARG"; } ]; + file_type = PYTHON_NUMPY; + field = [ { name = "gen_ens_prod_pyembed.py MET_PYTHON_INPUT_ARG"; } ]; In the event the user requires command line arguments to their Python script, they must be included alongside the file names separated by a delimiter. For example, the above Gen-Ens-Prod command with command line arguments for Python would look like: .. code-block:: - :caption: Gen-Ens-Prod Command Line with Python Args + :caption: Gen-Ens-Prod Command Line with Python Args - gen_ens_prod ens1.nc,arg1,arg2 ens2.nc,arg1,arg2 ens3.nc,arg1,arg2 ens4.nc,arg1,arg2 \ - -out ens_prod.nc -config GenEnsProd_config + gen_ens_prod -ens ens1.nc,arg1,arg2 ens2.nc,arg1,arg2 ens3.nc,arg1,arg2 ens4.nc,arg1,arg2 \ + -out ens_prod.nc -config GenEnsProd_config -In this case, the user's Python script will receive "ens1.nc,arg1,arg2" as a single command line argument for each execution of the Python script (i.e. 1 time per file). The user must parse this argument inside their Python script to obtain **arg1** and **arg2** as separate arguments. The list of input files and optionally, any command line arguments can be written to a single file (called **python_input_list** in the example below) that is substituted for the file names and command line arguments. ASCII file list elements are white-space separated (space-separated in the example below), as described in :numref:`ascii_file_lists`. For example: +In this case, the user's Python script will receive "ens1.nc,arg1,arg2" as a single command line argument for each execution of the Python script (i.e., 1 time per file). The user must parse this argument inside their Python script to obtain **arg1** and **arg2** as separate arguments. The list of input files and optionally, any command line arguments can be written to a single file (called **python_input_list** in the example below) that is substituted for the file names and command line arguments. ASCII file list elements are white-space separated (space-separated in the example below), as described in :numref:`ascii_file_lists`. For example: .. code-block:: - :caption: Gen-Ens-Prod File List + :caption: Gen-Ens-Prod File List - echo "file_list ens1.nc,arg1,arg2 ens2.nc,arg1,arg2 ens3.nc,arg1,arg2 ens4.nc,arg1,arg2" > python_input_list - gen_ens_prod python_input_list -out ens_prod.nc -config GenEnsProd_config + echo "file_list ens1.nc,arg1,arg2 ens2.nc,arg1,arg2 ens3.nc,arg1,arg2 ens4.nc,arg1,arg2" > python_input_list + gen_ens_prod -ens python_input_list -out ens_prod.nc -config GenEnsProd_config Finally, the above tools do not require data files to be present on a local disk. If the user wishes, their Python script can obtain data from other sources based upon only the command line arguments to their Python script. For example: .. code-block:: - :caption: Gen-Ens-Prod Python Args Only + :caption: Gen-Ens-Prod Python Args Only - gen_ens_prod 20230101,0 20230102,0 20230103,0 -out ens_prod.nc -confg GenEnsProd_config + gen_ens_prod -ens 20230101,0 20230102,0 20230103,0 -out ens_prod.nc -config GenEnsProd_config In the above command, each of the arguments "20230101,0", "20230102,0", and "20230103,0" are provided to the user's Python script in separate calls. Then, inside the Python script these arguments are used to construct a filename or query to a data server or other mechanism to return the desired data and format it the way MET expects inside the Python script, prior to calling Gen-Ens-Prod. @@ -416,28 +416,28 @@ Examples of Python Embedding for 2D Gridded Dataplanes **Grid-Stat with Python embedding for forecast and observations** .. code-block:: none - :caption: GridStat Command with Dual Python Embedding + :caption: GridStat Command with Dual Python Embedding - grid_stat 'PYTHON_NUMPY' 'PYTHON_NUMPY' GridStat_config -outdir /path/to/output + grid_stat 'PYTHON_NUMPY' 'PYTHON_NUMPY' GridStat_config -outdir /path/to/output .. code-block:: none - :caption: GridStat Config with Dual Python Embedding - - fcst = { - field = [ - { - name = "/path/to/fcst/python/script.py python_arg1 python_arg2"; - } - ]; - } - - obs = { - field = [ - { - name = "/path/to/obs/python/script.py python_arg1 python_arg2"; - } - ]; - } + :caption: GridStat Config with Dual Python Embedding + + fcst = { + field = [ + { + name = "/path/to/fcst/python/script.py python_arg1 python_arg2"; + } + ]; + } + + obs = { + field = [ + { + name = "/path/to/obs/python/script.py python_arg1 python_arg2"; + } + ]; + } .. _pyembed-point-obs-data: @@ -458,56 +458,56 @@ Python Script Requirements for Point Observations To provide the data that MET expects for point observations, the user is encouraged when designing their Python script to consider how to map their observations into the MET 11-column format. Then, the user can populate their observations into a Pandas DataFrame with the following column names and dtypes: .. list-table:: Point Observation DataFrame Columns and Dtypes - :widths: 5 5 10 - :header-rows: 1 - - * - column name - - data type (dtype) - - description - * - typ - - string - - Message Type - * - sid - - string - - Station ID - * - vld - - string - - Valid Time (YYYYMMDD_HHMMSS) - * - lat - - numeric - - Latitude (Degrees North) - * - lon - - numeric - - Longitude (Degrees East) - * - elv - - numeric - - Elevation (MSL) - * - var - - string - - Variable name (or GRIB code) - * - lvl - - numeric - - Level - * - hgt - - numeric - - Height (MSL or AGL) - * - qc - - string - - QC string - * - obs - - numeric - - Observation Value + :widths: 5 5 10 + :header-rows: 1 + + * - column name + - data type (dtype) + - description + * - typ + - string + - Message Type + * - sid + - string + - Station ID + * - vld + - string + - Valid Time (YYYYMMDD_HHMMSS) + * - lat + - numeric + - Latitude (Degrees North) + * - lon + - numeric + - Longitude (Degrees East) + * - elv + - numeric + - Elevation (MSL) + * - var + - string + - Variable name (or GRIB code) + * - lvl + - numeric + - Level + * - hgt + - numeric + - Height (MSL or AGL) + * - qc + - string + - QC string + * - obs + - numeric + - Observation Value To create the variable for MET, use the **.values** property of the Pandas DataFrame and the **.tolist()** method of the NumPy N-D Array. For example: .. code-block:: Python - :caption: Convert Pandas DataFrame to MET variable + :caption: Convert Pandas DataFrame to MET variable - # Pandas DataFrame - my_dataframe = pd.DataFrame() + # Pandas DataFrame + my_dataframe = pd.DataFrame() - # Convert to MET variable - point_data = my_dataframe.values.tolist() + # Convert to MET variable + point_data = my_dataframe.values.tolist() Running Python Embedding for Point Observations ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ @@ -515,20 +515,20 @@ Running Python Embedding for Point Observations The Point2Grid, Plot-Point-Obs, Ensemble-Stat, and Point-Stat tools support Python embedding for point observations. Python embedding for these tools can be invoked directly on the command line by replacing the input MET NetCDF point observation file name with the **full path** to the Python script and any arguments. The Python command must begin with the prefix **PYTHON_NUMPY=**. The full command should be enclosed in quotes to prevent embedded whitespace from causing parsing errors. An example of this is shown below for Plot-Point-Obs: .. code-block:: none - :caption: plot_point_obs with Python Embedding + :caption: plot_point_obs with Python Embedding - plot_point_obs \ - "PYTHON_NUMPY=scripts/python/examples/read_ascii_point.py data/sample_obs/ascii/sample_ascii_obs.txt" \ - output_image.ps + plot_point_obs \ + "PYTHON_NUMPY=scripts/python/examples/read_ascii_point.py data/sample_obs/ascii/sample_ascii_obs.txt" \ + output_image.ps The ASCII2NC tool also supports Python embedding, however invoking it varies slightly from other MET tools. For ASCII2NC, Python embedding is used by providing the "-format python" option on the command line. With this option, point observations may be passed as input. An example of this is shown below: .. code-block:: none - :caption: ascii2nc with Python Embedding + :caption: ascii2nc with Python Embedding - ascii2nc -format python \ - "scripts/python/examples/read_ascii_point.py data/sample_obs/ascii/sample_ascii_obs.txt" \ - sample_ascii_obs_python.nc + ascii2nc -format python \ + "scripts/python/examples/read_ascii_point.py data/sample_obs/ascii/sample_ascii_obs.txt" \ + sample_ascii_obs_python.nc Both of the above examples use the **read_ascii_point.py** example script which is included with the MET code. It reads ASCII data in MET's 11-column point observation format and stores it in a Pandas DataFrame to be read by the MET tools using Python embedding for point data. The **read_ascii_point.py** example script can be found in: @@ -542,20 +542,20 @@ Examples of Python Embedding for Point Observations **Point-Stat with Python embedding for forecast and observations** .. code-block:: none - :caption: PointStat Command with Dual Python Embedding + :caption: PointStat Command with Dual Python Embedding - point_stat 'PYTHON_NUMPY' 'PYTHON_NUMPY=/path/to/obs/python/script.py python_arg1 python_arg2' PointStat_config -outdir /path/to/output + point_stat 'PYTHON_NUMPY' 'PYTHON_NUMPY=/path/to/obs/python/script.py python_arg1 python_arg2' PointStat_config -outdir /path/to/output .. code-block:: none - :caption: PointStat Config with Dual Python Embedding - - fcst = { - field = [ - { - name = "/path/to/fcst/python/script.py python_arg1 python_arg2"; - } - ]; - } + :caption: PointStat Config with Dual Python Embedding + + fcst = { + field = [ + { + name = "/path/to/fcst/python/script.py python_arg1 python_arg2"; + } + ]; + } .. _pyembed-mpr-data: @@ -568,7 +568,7 @@ The MET Pair-Stat tool also supports Python embedding of matched pair (MPR) data .. note:: - While Stat-Analysis can read all STAT line types through Python embedding, Pair-Stat only reads the MPR line type. Note that the MET statistics tools write all output line types to a STAT file, but can also be configured to write each line type to separate text (TXT) files. The example below reads data from an MPR text file generated by Point-Stat where each line has the same number of columns. It will not work for STAT files, in general, where the number of columns vary by line type. + While Stat-Analysis can read all STAT line types through Python embedding, Pair-Stat only reads the MPR line type. Note that the MET statistics tools write all output line types to a STAT file, but can also be configured to write each line type to separate text (TXT) files. The example below reads data from an MPR text file generated by Point-Stat where each line has the same number of columns. It will not work for STAT files, in general, where the number of columns varies by line type. Python Script Requirements for MPR Data ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ @@ -577,28 +577,28 @@ Python Script Requirements for MPR Data 2. The **mpr_data** variable must be a Python list representation of a NumPy N-D Array created from a Pandas DataFrame -3. The **met_data** variable must have data in **exactly** 36 columns for MPR data, corresponding to the summation of the :ref:`common STAT output` and the :ref:`MPR line type output`. +3. The **mpr_data** variable must have data in **exactly** 36 columns for MPR data, corresponding to the summation of the :ref:`common STAT output` and the :ref:`MPR line type output`. If a user does not have an existing MPR line type file created by the MET tools, they will need to map their data into the 36 columns expected by Stat-Analysis for the MPR line type data. If a user already has MPR line type files, the most direct way for a user to read MPR line type data is to model their Python script after the sample **read_ascii_mpr.py** script. Sample code is included here for convenience: .. code-block:: Python - :caption: Reading MPR line types with Pandas + :caption: Reading MPR line types with Pandas - # Open the MPR line type file - mpr_dataframe = pd.read_csv(input_mpr_file,\ - header=None,\ - delim_whitespace=True,\ - keep_default_na=False,\ - skiprows=1,\ - usecols=range(1,36),\ - dtype=str) + # Open the MPR line type file + mpr_dataframe = pd.read_csv(input_mpr_file,\ + header=None,\ + delim_whitespace=True,\ + keep_default_na=False,\ + skiprows=1,\ + usecols=range(1,36),\ + dtype=str) - # Convert to the variable MET expects - mpr_data = mpr_dataframe.values.tolist() + # Convert to the variable MET expects + mpr_data = mpr_dataframe.values.tolist() .. note:: - If reading non-MPR STAT line types as input, the example above should be modified based on the number of columns for the input line type. + If reading non-MPR STAT line types as input, the example above should be modified based on the number of columns for the input line type. Running Python Embedding for MPR Data ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ @@ -606,12 +606,12 @@ Running Python Embedding for MPR Data Stat-Analysis can be run using the **-lookin python** command line option: .. code-block:: none - :caption: Stat-Analysis with Python Embedding of MPR Data + :caption: Stat-Analysis with Python Embedding of MPR Data - stat_analysis \ - -lookin python scripts/python/examples/read_ascii_mpr.py point_stat_mpr.txt \ - -job aggregate_stat -line_type MPR -out_line_type CNT \ - -by FCST_VAR,FCST_LEV + stat_analysis \ + -lookin python scripts/python/examples/read_ascii_mpr.py point_stat_mpr.txt \ + -job aggregate_stat -line_type MPR -out_line_type CNT \ + -by FCST_VAR,FCST_LEV In this example, rather than passing the MPR output lines from Point-Stat directly into Stat-Analysis (which is the typical approach), the **read_ascii_mpr.py** Python embedding script reads that file and passes the data to Stat-Analysis. The aggregate_stat job is defined on the command line and CNT statistics are derived from the MPR input data. Separate CNT statistics are computed for each unique combination of FCST_VAR and FCST_LEV present in the input. @@ -629,8 +629,8 @@ MET comes with a Python package that provides core functionality for the Python To utilize the MET Python package **standalone** when NOT using it with Python embedding, users must add the following to their **PYTHONPATH** environment variable: .. code-block:: - :caption: MET Python Module PYTHONPATH + :caption: MET Python Module PYTHONPATH - export PYTHONPATH={MET_INSTALL_DIR}/share/met/python + export PYTHONPATH={MET_INSTALL_DIR}/share/met/python where {MET_INSTALL_DIR} is the top level directory where MET is installed, for example **/usr/local/met**. diff --git a/docs/Users_Guide/appendixG.rst b/docs/Users_Guide/appendixG.rst index 90f653c60b..5b5a07f342 100644 --- a/docs/Users_Guide/appendixG.rst +++ b/docs/Users_Guide/appendixG.rst @@ -74,9 +74,9 @@ FBAR and OBAR are the average values of the forecast and observed wind speed. .. math:: - \text{FBAR} = \frac{1}{N} \sum_i s_{fi} + \text{FBAR} = \frac{1}{N} \sum_i s_{fi} - \text{OBAR} = {1 \over N} \sum_i s_{oi} + \text{OBAR} = {1 \over N} \sum_i s_{oi} _________________________ @@ -84,19 +84,19 @@ FS_RMS and OS_RMS are the root-mean-square values of the forecast and observed w .. only:: latex - .. math:: + .. math:: - \text{FS\_RMS} = [ \frac{1}{N} \sum_i s_{fi}^2]^{1/2} + \text{FS\_RMS} = [ \frac{1}{N} \sum_i s_{fi}^2]^{1/2} - \text{OS\_RMS} = [\frac{1}{N} \sum_i s_{oi}^2]^{1/2} + \text{OS\_RMS} = [\frac{1}{N} \sum_i s_{oi}^2]^{1/2} .. only:: html - .. math:: + .. math:: - \text{FS_RMS} = [ \frac{1}{N} \sum_i s_{fi}^2]^{1/2} + \text{FS_RMS} = [ \frac{1}{N} \sum_i s_{fi}^2]^{1/2} - \text{OS_RMS} = [\frac{1}{N} \sum_i s_{oi}^2]^{1/2} + \text{OS_RMS} = [\frac{1}{N} \sum_i s_{oi}^2]^{1/2} ___________________________ @@ -104,9 +104,9 @@ MSVE and RMSVE are, respectively, the mean squared, and root mean squared, lengt .. math:: - \text{MSVE} = \frac{1}{N} \sum_i | \mathbf{F}_i - \mathbf{O}_i|^2 + \text{MSVE} = \frac{1}{N} \sum_i | \mathbf{F}_i - \mathbf{O}_i|^2 - \text{RMSVE} = \sqrt{MSVE} + \text{RMSVE} = \sqrt{MSVE} ____________________________ @@ -114,7 +114,7 @@ FSTDEV and OSTDEV are the standard deviations of the forecast and observed wind .. math:: \text{FSTDEV } = \frac{1}{N} \sum_i (s_{fi} - \text{FBAR})^2 = \frac{1}{N} \sum_i s_{fi}^2 - \text{FBAR}^2 - \text{OSTDEV } = \frac{1}{N} \sum_i (s_{oi} - \text{OBAR})^2 = \frac{1}{N} \sum_i s_{oi}^2 - \text{OBAR}^2 + \text{OSTDEV } = \frac{1}{N} \sum_i (s_{oi} - \text{OBAR})^2 = \frac{1}{N} \sum_i s_{oi}^2 - \text{OBAR}^2 ___________________________ @@ -122,47 +122,47 @@ FDIR and ODIR are the direction (angle) of :math:`\mathbf{F}_a \text{ and } \mat .. math:: \text{FDIR } = \text{ direction angle of } \mathbf{F}_a - \text{ODIR} = \text{ direction angle of } \mathbf{O}_a + \text{ODIR} = \text{ direction angle of } \mathbf{O}_a ________________________ -FBAR_SPEED and OBAR_SPEED are the lengths of the average forecast and observed wind vectors. Note that this is *not* the same as the average forecast and observed wind speeds (*ie.,* the length of an average vector :math:`\neq` the average length of the vector). +FBAR_SPEED and OBAR_SPEED are the lengths of the average forecast and observed wind vectors. Note that this is *not* the same as the average forecast and observed wind speeds (*i.e.,* the length of an average vector :math:`\neq` the average length of the vector). .. only:: latex - .. math:: + .. math:: - \text{FBAR\_SPEED } = | \mathbf{F}_a | + \text{FBAR\_SPEED } = | \mathbf{F}_a | - \text{OBAR\_SPEED } = | \mathbf{O}_a | + \text{OBAR\_SPEED } = | \mathbf{O}_a | .. only:: html - .. math:: + .. math:: - \text{FBAR_SPEED } = | \mathbf{F}_a | + \text{FBAR_SPEED } = | \mathbf{F}_a | - \text{OBAR_SPEED } = | \mathbf{O}_a | + \text{OBAR_SPEED } = | \mathbf{O}_a | ________________________ -VDIFF_SPEED is the length (*ie. speed*) of the vector difference between the average forecast and average observed wind vectors. +VDIFF_SPEED is the length (*i.e., speed*) of the vector difference between the average forecast and average observed wind vectors. .. only:: latex - .. math:: \text{VDIFF\_SPEED } = | \mathbf{F}_a - \mathbf{O}_a | + .. math:: \text{VDIFF\_SPEED } = | \mathbf{F}_a - \mathbf{O}_a | .. only:: html - .. math:: \text{VDIFF_SPEED } = | \mathbf{F}_a - \mathbf{O}_a | + .. math:: \text{VDIFF_SPEED } = | \mathbf{F}_a - \mathbf{O}_a | .. only:: latex - Note that this is *not* the same as the difference in lengths (speeds) of the average forecast and observed wind vectors. That quantity is called SPEED_ERR (see below). There is a relationship between these two statistics however: using some of the results obtained in the introduction to this appendix, we can say that :math:`| | \mathbf{F}_a | - | \mathbf{O}_a | | \leq | \mathbf{F}_a - \mathbf{O}_a |` or, equivalently, that :math:`\vert \text{SPEED\_ERR } \vert \leq \text{VDIFF\_SPEED. }` + Note that this is *not* the same as the difference in lengths (speeds) of the average forecast and observed wind vectors. That quantity is called SPEED_ERR (see below). There is a relationship between these two statistics however: using some of the results obtained in the introduction to this appendix, we can say that :math:`| | \mathbf{F}_a | - | \mathbf{O}_a | | \leq | \mathbf{F}_a - \mathbf{O}_a |` or, equivalently, that :math:`\vert \text{SPEED\_ERR } \vert \leq \text{VDIFF\_SPEED. }` .. only:: html - Note that this is *not* the same as the difference in lengths (speeds) of the average forecast and observed wind vectors. That quantity is called SPEED_ERR (see below). There is a relationship between these two statistics however: using some of the results obtained in the introduction to this appendix, we can say that :math:`| | \mathbf{F}_a | - | \mathbf{O}_a | | \leq | \mathbf{F}_a - \mathbf{O}_a |` or, equivalently, that :math:`\vert \text{SPEED_ERR } \vert \leq \text{VDIFF_SPEED. }` + Note that this is *not* the same as the difference in lengths (speeds) of the average forecast and observed wind vectors. That quantity is called SPEED_ERR (see below). There is a relationship between these two statistics however: using some of the results obtained in the introduction to this appendix, we can say that :math:`| | \mathbf{F}_a | - | \mathbf{O}_a | | \leq | \mathbf{F}_a - \mathbf{O}_a |` or, equivalently, that :math:`\vert \text{SPEED_ERR } \vert \leq \text{VDIFF_SPEED. }` _________________________ @@ -170,11 +170,11 @@ VDIFF_DIR is the direction of the vector difference of the average forecast and .. only:: latex - .. math:: \text{VDIFF\_DIR } = \text{ direction of } (\mathbf{F}_a - \mathbf{O}_a) + .. math:: \text{VDIFF\_DIR } = \text{ direction of } (\mathbf{F}_a - \mathbf{O}_a) .. only:: html - .. math:: \text{VDIFF_DIR } = \text{ direction of } (\mathbf{F}_a - \mathbf{O}_a) + .. math:: \text{VDIFF_DIR } = \text{ direction of } (\mathbf{F}_a - \mathbf{O}_a) _________________________ @@ -182,11 +182,11 @@ SPEED_ERR is the difference in the lengths (speeds) of the average forecast and .. only:: latex - .. math:: \text{SPEED\_ERR } = | \mathbf{F}_a | - | \mathbf{O}_a | = \text{ FBAR\_SPEED } - \text{ OBAR\_SPEED } + .. math:: \text{SPEED\_ERR } = | \mathbf{F}_a | - | \mathbf{O}_a | = \text{ FBAR\_SPEED } - \text{ OBAR\_SPEED } .. only:: html - .. math:: \text{SPEED_ERR } = | \mathbf{F}_a | - | \mathbf{O}_a | = \text{ FBAR_SPEED } - \text{ OBAR_SPEED } + .. math:: \text{SPEED_ERR } = | \mathbf{F}_a | - | \mathbf{O}_a | = \text{ FBAR_SPEED } - \text{ OBAR_SPEED } ___________________________ @@ -194,11 +194,11 @@ SPEED_ABSERR is the absolute value of SPEED_ERR. Note that we have SPEED_ABSERR .. only:: latex - .. math:: \text{SPEED\_ABSERR } = \vert \text{SPEED\_ERR } \vert + .. math:: \text{SPEED\_ABSERR } = \vert \text{SPEED\_ERR } \vert .. only:: html - .. math:: \text{SPEED_ABSERR } = \vert \text{SPEED_ERR } \vert + .. math:: \text{SPEED_ABSERR } = \vert \text{SPEED_ERR } \vert __________________________ @@ -206,11 +206,11 @@ DIR_ERR is the signed angle between the directions of the average forecast and a .. only:: latex - .. math:: \text{DIR\_ERR } = \text{ direction between } N(\mathbf{F}_a) \text{ and } N(\mathbf{O}_a) + .. math:: \text{DIR\_ERR } = \text{ direction between } N(\mathbf{F}_a) \text{ and } N(\mathbf{O}_a) .. only:: html - .. math:: \text{DIR_ERR } = \text{ direction between } N(\mathbf{F}_a) \text{ and } N(\mathbf{O}_a) + .. math:: \text{DIR_ERR } = \text{ direction between } N(\mathbf{F}_a) \text{ and } N(\mathbf{O}_a) __________________________ @@ -218,11 +218,11 @@ DIR_ABSERR is the absolute value of DIR_ERR. In other words, it's an unsigned an .. only:: latex - .. math:: \text{DIR\_ABSERR } = \vert \text{DIR\_ERR } \vert + .. math:: \text{DIR\_ABSERR } = \vert \text{DIR\_ERR } \vert .. only:: html - .. math:: \text{DIR_ABSERR } = \vert \text{DIR_ERR } \vert + .. math:: \text{DIR_ABSERR } = \vert \text{DIR_ERR } \vert __________________________ @@ -232,7 +232,7 @@ For each point, the directed angle difference in degrees is computed between the .. math:: - N(\mathbf{F}_i) - N(\mathbf{O}_i) \in (-180, 180] + N(\mathbf{F}_i) - N(\mathbf{O}_i) \in (-180, 180] Note however that the direction of the zero vector is undefined. Points for which the forecast or observed wind direction is undefined are excluded from the analysis and result in a warning message being printed. The "wind_thresh" and "wind_logic" configuration options, described in :numref:`config_options`, can be used to filter the wind vectors down to a subset that meet the specified wind speed threshold. @@ -242,15 +242,15 @@ DIR_ME is the average of the signed difference between the forecast and observed .. only:: latex - .. math:: + .. math:: - \text{DIR\_ME} = \frac{1}{N} \sum_i (N(\mathbf{F}_i) - N(\mathbf{O}_i)) + \text{DIR\_ME} = \frac{1}{N} \sum_i (N(\mathbf{F}_i) - N(\mathbf{O}_i)) .. only:: html - .. math:: + .. math:: - \text{DIR_ME} = \frac{1}{N} \sum_i (N(\mathbf{F}_i) - N(\mathbf{O}_i)) + \text{DIR_ME} = \frac{1}{N} \sum_i (N(\mathbf{F}_i) - N(\mathbf{O}_i)) __________________________ @@ -258,15 +258,15 @@ DIR_MAE is the average of the absolute value of the difference between the forec .. only:: latex - .. math:: + .. math:: - \text{DIR\_MAE} = \frac{1}{N} \sum_i | N(\mathbf{F}_i) - N(\mathbf{O}_i) | + \text{DIR\_MAE} = \frac{1}{N} \sum_i | N(\mathbf{F}_i) - N(\mathbf{O}_i) | .. only:: html - .. math:: + .. math:: - \text{DIR_MAE} = \frac{1}{N} \sum_i | N(\mathbf{F}_i) - N(\mathbf{O}_i) | + \text{DIR_MAE} = \frac{1}{N} \sum_i | N(\mathbf{F}_i) - N(\mathbf{O}_i) | __________________________ @@ -274,15 +274,15 @@ DIR_MSE is the average of the squared difference between the forecast and observ .. only:: latex - .. math:: + .. math:: - \text{DIR\_MSE} = \frac{1}{N} \sum_i (N(\mathbf{F}_i) - N(\mathbf{O}_i))^2 + \text{DIR\_MSE} = \frac{1}{N} \sum_i (N(\mathbf{F}_i) - N(\mathbf{O}_i))^2 .. only:: html - .. math:: + .. math:: - \text{DIR_MSE} = \frac{1}{N} \sum_i (N(\mathbf{F}_i) - N(\mathbf{O}_i))^2 + \text{DIR_MSE} = \frac{1}{N} \sum_i (N(\mathbf{F}_i) - N(\mathbf{O}_i))^2 __________________________ @@ -290,12 +290,12 @@ DIR_RMSE is the square root of the average squared difference between the foreca .. only:: latex - .. math:: + .. math:: - \text{DIR\_RMSE} = \sqrt{DIR\_MSE} + \text{DIR\_RMSE} = \sqrt{DIR\_MSE} .. only:: html - .. math:: + .. math:: - \text{DIR_RMSE} = \sqrt{DIR\_MSE} + \text{DIR_RMSE} = \sqrt{DIR\_MSE} diff --git a/docs/Users_Guide/appendixH.rst b/docs/Users_Guide/appendixH.rst index 52a33ea903..a42152d7a3 100644 --- a/docs/Users_Guide/appendixH.rst +++ b/docs/Users_Guide/appendixH.rst @@ -66,7 +66,7 @@ If the UGRID dataset is supported by MET, then this item is set to the string id ugrid_max_distance_km --------------------- -For PointStat, this is the maximum allowable distance in kilometers for a point observation to be matched to the nearest UGRID cell center. If left unspecified, each point observation will be matched to the closest UGRID cell, regardless of how far apart they are. If no UGRID cells are within the specified distance, the observation will be not used for verification. For GridStat, this is the distance from each forecast grid point to search for observation grid cells to include when interpolating to the forecast grid point locations. Currently, only the nearest point observation or gridded observation cell is used. See `Unstructured Grid Limitations`_ for additional details about interpolation when using GridStat. +For PointStat, this is the maximum allowable distance in kilometers for a point observation to be matched to the nearest UGRID cell center. If left unspecified, each point observation will be matched to the closest UGRID cell, regardless of how far apart they are. If no UGRID cells are within the specified distance, the observation will not be used for verification. For GridStat, this is the distance from each forecast grid point to search for observation grid cells to include when interpolating to the forecast grid point locations. Currently, only the nearest point observation or gridded observation cell is used. See `Unstructured Grid Limitations`_ for additional details about interpolation when using GridStat. ugrid_coordinates_file ---------------------- diff --git a/docs/Users_Guide/config_options.rst b/docs/Users_Guide/config_options.rst index 8198e17228..9c9b5d4832 100644 --- a/docs/Users_Guide/config_options.rst +++ b/docs/Users_Guide/config_options.rst @@ -48,7 +48,7 @@ The configuration file language supports the following data types: * Threshold: - * A threshold type (<, <=, ==, !-, >=, or >) followed by a numeric value. + * A threshold type (<, <=, ==, !=, >=, or >) followed by a numeric value. * The threshold type may also be specified using two letter abbreviations (lt, le, eq, ne, ge, gt). @@ -88,26 +88,26 @@ The configuration file language supports the following data types: * The following percentile threshold types are supported: * SFP for a percentile of the sample forecast values. - e.g. ">SFP33.3" means greater than the 33.3-rd forecast percentile. + e.g., ">SFP33.3" means greater than the 33.3-rd forecast percentile. * SOP for a percentile of the sample observation values. - e.g. ">SOP75" means greater than the 75-th observation percentile. + e.g., ">SOP75" means greater than the 75-th observation percentile. * SFCP for a percentile of the sample forecast climatology values. - e.g. ">SFCP90" means greater than the 90-th forecast climatology + e.g., ">SFCP90" means greater than the 90-th forecast climatology percentile. * SOCP for a percentile of the sample observation climatology values. - e.g. ">SOCP90" means greater than the 90-th observation climatology + e.g., ">SOCP90" means greater than the 90-th observation climatology percentile. For backward compatibility, the "SCP" threshold type is processed the same as "SOCP". * USP for a user-specified percentile threshold. - e.g. "5.0" @@ -122,7 +122,7 @@ The configuration file language supports the following data types: conditional (cnt_thresh), or wind speed (wind_thresh) thresholds can be defined relative to the climatological distribution at each point. Therefore, the actual numeric threshold applied can change for each point. - e.g. ">FCDP50" means greater than the 50-th percentile of the + e.g., ">FCDP50" means greater than the 50-th percentile of the climatological distribution for each point. * OCDP for observation climatological distribution percentile thresholds. @@ -170,13 +170,13 @@ The configuration file language supports the following data types: * Piecewise-Linear Function (currently used only by MODE): - * A list of (x, y) points enclosed in parenthesis (). + * A list of (x, y) points enclosed in parentheses (). * The (x, y) points are *NOT* separated by commas. * User-defined function of a single variable: - * Left side is a function name followed by variable name in parenthesis. + * Left side is a function name followed by variable name in parentheses. * Right side is an equation which includes basic math functions (+,-,*,/), built-in functions (listed below), or other user-defined functions. @@ -205,13 +205,13 @@ and *scripts/config*. When you pass a configuration file to a MET tool, the tool actually parses up to four different configuration files in the following order: - 1. Reads *share/met/config/ConfigConstants* to define constants. + 1. Reads *share/met/config/ConfigConstants* to define constants. - 2. If the tool produces PostScript output, it reads *share/met/config/ConfigMapData* to define the map data to be plotted. + 2. If the tool produces PostScript output, it reads *share/met/config/ConfigMapData* to define the map data to be plotted. - 3. Reads the default configuration file for the tool from *share/met/config*. + 3. Reads the default configuration file for the tool from *share/met/config*. - 4. Reads the user-specified configuration file from the command line. + 4. Reads the user-specified configuration file from the command line. Many of the entries from step (3) are overwritten by the user-specified entries from step (4). Therefore, the configuration file you pass in on the command @@ -274,12 +274,12 @@ The MET_AIRNOW_STATIONS environment variable can be used to specify a file that will override the default file. If set, it should be the full path to the file. The default table can be found in the installed *share/met/table_files/airnow_monitoring_site_locations_v2.dat*. This file contains -ascii column data that allows lookups of latitude, longitude, and elevation for all +ASCII column data that allows lookups of latitude, longitude, and elevation for all AirNow stations based on stationId and/or AqSid. Additional information and updated site locations can be found at the `EPA AirNow website `_. While some monitoring stations are -permanent, others are temporary, and theirs locations can change. When running the +permanent, others are temporary, and their locations can change. When running the ASCII2NC tool with the :code:`-format airnowhourly` option, users should `download `_ the **Monitoring_Site_Locations_V2.dat** data file for the date being processed and set the MET_AIRNOW_STATIONS environment @@ -291,7 +291,7 @@ MET_NDBC_STATIONS ----------------- The MET_NDBC_STATIONS environment variable can be used to specify a file that -will override the default file. If set it should be a full path to the file. +will override the default file. If set, it should be a full path to the file. The default table can be found in the installed *share/met/table_files/ndbc_stations.xml*. This file contains XML content for all stations that allows lookups of latitude, longitude, @@ -299,7 +299,7 @@ and, in some cases, elevation for all stations based on stationId. This set of stations comes from 2 online sources: the `active stations website `_ -and the `complete stations website `_. +and the `complete stations website `_. As these lists can change as a function of time, a script can be run to pull down the contents of both websites and merge any changes with the existing stations @@ -317,7 +317,7 @@ To run this utility: Usage: build_ndbc_stations_from_web.py [options] Options: -h, --help show this help message and exit - -d, --diagnostic Rerun using downlaoded files, skipping download step (optional, default: False) + -d, --diagnostic Rerun using downloaded files, skipping download step (optional, default: False) -p, --prune Prune files that are no longer online (optional, default: False) -o OUT_FILE, --out=OUT_FILE Save the text into the named file (optional, default: merged.txt) @@ -398,9 +398,9 @@ GRIB1 table files begin with "grib1" prefix and end with a ".txt" suffix. The first line of the file must contain GRIB1. The following lines consist of 4 integers followed by 3 strings: -| Column 1: GRIB code (e.g. 11 for temperature) +| Column 1: GRIB code (e.g., 11 for temperature) | Column 2: parameter table version number -| Column 3: center id (e.g. 07 for US Weather Service- National Met. Center) +| Column 3: center id (e.g., 07 for US Weather Service- National Met. Center) | Column 4: subcenter id | Column 5: variable name | Column 6: variable description @@ -456,7 +456,7 @@ Further expanded use of parallelism is planned for future versions of MET. Due to the broad application of OpenMP, nearly all MET applications benefit from it. However, initializing OpenMP threads does incur some overhead cost. Typically the runtime benefit dramatically outweighs the setup cost. Generally, -more threads produces faster runtimes, but that is not always the case. The +more threads produce faster runtimes, but that is not always the case. The optimal number of threads for any single run of a MET tool is data dependent. Setting the number of threads @@ -503,7 +503,7 @@ observed in practice, however. A lower thread count is appropriate when time-to-solution is not so critical, because cores remain idle when the code is not inside a parallel region. Fewer -threads typically means better resource utilization. +threads typically mean better resource utilization. Thread Binding ^^^^^^^^^^^^^^ @@ -538,7 +538,7 @@ and debugging to keep them for further inspection. Setting this environment vari to a value of :code:`yes` or :code:`true` instructs the MET tools to retain temporary files instead of deleting them. -Note that doing so may fill up the temporary directory. It is the responsiblity of +Note that doing so may fill up the temporary directory. It is the responsibility of the user to monitor the temporary directory usage and remove temporary files that are no longer needed. @@ -576,7 +576,7 @@ MET_USE_WRF_SUBGRID The MET_USE_WRF_SUBGRID environment variable controls how the grid information is read from WRF files. -Some WRF files contain fields that are on a subgrid, which contain more grid points +Some WRF files contain fields that are on a subgrid, which contains more grid points and require a computation to determine the d_km value. MET reads the grid information from a file before any fields are read and assumes that there is one grid definition per file. @@ -623,7 +623,7 @@ The "nc_compression" entry in ConfigConstants defines the compression level for the NetCDF variables. Setting this option in the config file of one of the tools overrides the default value set in ConfigConstants. The environment variable MET_NC_COMPRESS overrides the compression level -from configuration file. The command line argument "-compress n" for some +from the configuration file. The command line argument "-compress n" for some tools overrides it. The range is 0 to 9. @@ -659,7 +659,7 @@ The "tmp_dir" entry in ConfigConstants defines the directory for the temporary files. The directory must exist and be writable. The environment variable MET_TMP_DIR overrides the default value at the configuration file. Some tools override the temporary directory by the command line argument -"-tmp_dir ". +"-tmp_dir ". .. code-block:: none @@ -700,7 +700,7 @@ default values found in the "data/config/ConfigConstants" file: wind_direction_field_name = "WDIR,DD"; Each is a comma-separated list of wind variable names to be searched. Users can -explicity set these options to configure what data should be used in the wind +explicitly set these options to configure what data should be used in the wind derivation and rotation logic. message_type_group_map @@ -751,9 +751,9 @@ The "obtype_as_group_val_flag" entry is a boolean that controls how the OBTYPE header column is populated for message type groups defined in "message_type_group_map". If set to TRUE and when writing matched pair line types (MPR, SEEPS_MPR, and ORANK), write OBTYPE as the group map -*value*, i.e. the input message type for each individual observation. +*value*, i.e., the input message type for each individual observation. If set to FALSE (default) and for all other line types, write OBTYPE -as the group map key, i.e. the name of the message type group. +as the group map key, i.e., the name of the message type group. For example, if FALSE, write the OBTYPE column in the MPR line type as the "ANYAIR" message type group name. If TRUE, write OBTYPE as "AIRCAR" @@ -761,7 +761,7 @@ or "AIRCFT", based on the input message type of each point observation. .. code-block:: none - obtyp_as_group_val_flag = FALSE; + obtype_as_group_val_flag = FALSE; message_type_map ---------------- @@ -770,7 +770,7 @@ The "message_type_map" entry is an array of dictionaries, each containing a "key" string and "val" string. This defines a mapping of input strings to output message types. This mapping is applied in ASCII2NC when converting input little_r report types to output message types. This mapping -is also supported in PBN2NC as a way of renaming input PREPBUFR message +is also supported in PB2NC as a way of renaming input PREPBUFR message types. .. code-block:: none @@ -794,7 +794,7 @@ The "model" entry specifies a name for the model being verified. This name is written to the MODEL column of the ASCII output generated. If you're verifying multiple models, you should choose descriptive model names (no whitespace) to distinguish between their output. -e.g. model = "GFS"; +e.g., model = "GFS"; .. code-block:: none @@ -810,7 +810,7 @@ entry or simply once at the top level of the configuration file. If you're verifying the same field multiple times with different quality control flags, you should choose description strings (no whitespace) to distinguish between their output. -e.g. desc = "QC_9"; +e.g., desc = "QC_9"; .. code-block:: none @@ -949,7 +949,7 @@ smoothing. The default is 120. Ignored if not Gaussian method. .. note:: The "gaussian_dx" and "gaussian_radius" settings must be in the same - units, such as kilometers or degress. Their ratio + units, such as kilometers or degrees. Their ratio (sigma = gaussian_radius / gaussian_dx) determines the Gaussian weighting function. @@ -1004,7 +1004,7 @@ The "level" entry specifies level information for the field. Setting field.prob """""""""" The "prob" entry in the forecast dictionary defines probability -information. It may either be set as a boolean (i.e. TRUE or FALSE) +information. It may either be set as a boolean (i.e., TRUE or FALSE) or as a dictionary defining probabilistic field information. When set as a boolean to TRUE, it indicates that the "fcst.field" data @@ -1044,10 +1044,10 @@ data, one could configure the Grid-Stat or Point-Stat tools as follows: The example above selects two probabilistic fields. In both, "name" is set to "PROB", the GRIB abbreviation for probabilities. The "level" -entry defines the level information (i.e. "A24" for a 24-hour +entry defines the level information (i.e., "A24" for a 24-hour accumulation and "P850" for 850mb). The "prob" dictionary defines the event for which the probability is defined. The "thresh_lo" -(i.e. APCP > 2.54) and/or "thresh_hi" (i.e. TMP < 273) entries are +(i.e., APCP > 2.54) and/or "thresh_hi" (i.e., TMP < 273) entries are used to define the event threshold(s). Probability fields should contain values in the range @@ -1097,7 +1097,7 @@ The "convert" entry is a user-defined function of a single variable for processing input data values. Any input values that are not bad data are replaced by the value of this function. The convert function is applied prior to regridding or thresholding. This function may -include any of the built-in math functions (e.g. sqrt, log10) +include any of the built-in math functions (e.g., sqrt, log10) described above. Several standard unit conversion functions are already defined in *data/config/ConfigConstants*. @@ -1141,7 +1141,7 @@ field.mpr_column and field.mpr_thresh The "mpr_column" and "mpr_thresh" entries are arrays of strings and thresholds to specify which matched pairs should be included in the statistics. These options apply to the Point-Stat and Grid-Stat tools. -They are parsed seperately for each "obs.field" array entry. +They are parsed separately for each "obs.field" array entry. The "mpr_column" strings specify MPR column names (FCST, OBS, CLIMO_MEAN, CLIMO_STDEV, or CLIMO_CDF), differences of columns (FCST-OBS), or the absolute value of those differences (ABS(FCST-OBS)). @@ -1215,31 +1215,31 @@ The "file_type" entry specifies the input gridded data file type rather than letting the code determine it. MET determines the file type by checking for known suffixes and examining the file contents. Use this option to override the code's choice. The valid file_type values are -listed the "data/config/ConfigConstants" file and are described below. +listed in the "data/config/ConfigConstants" file and are described below. This entry should be defined within the "fcst" and/or "obs" dictionaries. For example: - .. code-block:: none +.. code-block:: none - fcst = { - file_type = GRIB1; GRIB version 1 - file_type = GRIB2; GRIB version 2 - file_type = NETCDF_MET; NetCDF created by another MET tool - file_type = NETCDF_WRF; NetCDF WRF output. - file_type = NETCDF_PINT; NetCDF created by running the p_interp - or wrf_interp utility on WRF output. - May be used to read unstaggered raw WRF - NetCDF output at the surface or a - single model level. - file_type = NETCDF_NCCF; NetCDF following the Climate Forecast - (CF) convention. - file_type = NETCDF_UGRID; NetCDF containing data on an - unstructured grid. - file_type = PYTHON_NUMPY; Run a Python script to load data into - a NumPy array. - file_type = PYTHON_XARRAY; Run a Python script to load data into - an xarray object. - } + fcst = { + file_type = GRIB1; GRIB version 1 + file_type = GRIB2; GRIB version 2 + file_type = NETCDF_MET; NetCDF created by another MET tool + file_type = NETCDF_WRF; NetCDF WRF output. + file_type = NETCDF_PINT; NetCDF created by running the p_interp + or wrf_interp utility on WRF output. + May be used to read unstaggered raw WRF + NetCDF output at the surface or a + single model level. + file_type = NETCDF_NCCF; NetCDF following the Climate Forecast + (CF) convention. + file_type = NETCDF_UGRID; NetCDF containing data on an + unstructured grid. + file_type = PYTHON_NUMPY; Run a Python script to load data into + a NumPy array. + file_type = PYTHON_XARRAY; Run a Python script to load data into + an xarray object. + } wind_thresh ^^^^^^^^^^^ @@ -1329,7 +1329,7 @@ GRIB1 and GRIB2 * The "GRIB_lvl_typ" entry is an integer specifying the level type. * The "GRIB_lvl_val1" and "GRIB_lvl_val2" entries are floats specifying - the first and second level values. + the first and second level values. * The "GRIB_ens" entry is a string specifying NCEP's usage of the extended PDS for ensembles. Set to "hi_res_ctl", "low_res_ctl", @@ -1389,8 +1389,8 @@ GRIB1 and GRIB2 templates 4.46 and 4.48. * The GRIB2_aerosol_size_lower and "GRIB2_aerosol_size_upper" are doubles - specifying the endpoints of the aerosol size interval. These applies only - to GRIB2 product defintion templates 4.46 and 4.48. + specifying the endpoints of the aerosol size interval. These apply only + to GRIB2 product definition templates 4.46 and 4.48. * The GRIB2_ipdtmpl_index and GRIB2_ipdtmpl_val entries are arrays of integers which specify the product description template values to @@ -1466,9 +1466,9 @@ Using PYTHON_NUMPY or PYTHON_XARRAY: .. code-block:: none - field = [ - { name = "read_ascii_numpy.py data/python/fcst.txt FCST"; } - ]; + field = [ + { name = "read_ascii_numpy.py data/python/fcst.txt FCST"; } + ]; Option 2: @@ -1530,48 +1530,48 @@ length of the "fcst.field" array. For example: .. code-block:: none - obs = fcst; + obs = fcst; or .. code-block:: none - fcst = { - censor_thresh = []; - censor_val = []; - cnt_thresh = [ NA ]; - cnt_logic = UNION; - wind_thresh = [ NA ]; - wind_logic = UNION; - - field = [ - { - name = "PWAT"; - level = [ "L0" ]; - cat_thresh = [ >2.5 ]; - } - ]; - } - + fcst = { + censor_thresh = []; + censor_val = []; + cnt_thresh = [ NA ]; + cnt_logic = UNION; + wind_thresh = [ NA ]; + wind_logic = UNION; + + field = [ + { + name = "PWAT"; + level = [ "L0" ]; + cat_thresh = [ >2.5 ]; + } + ]; + } - obs = { - censor_thresh = []; - censor_val = []; - mpr_column = []; - mpr_thresh = []; - cnt_thresh = [ NA ]; - cnt_logic = UNION; - wind_thresh = [ NA ]; - wind_logic = UNION; - field = [ - { - name = "IWV"; - level = [ "L0" ]; - cat_thresh = [ >25.0 ]; - } - ]; - } + obs = { + censor_thresh = []; + censor_val = []; + mpr_column = []; + mpr_thresh = []; + cnt_thresh = [ NA ]; + cnt_logic = UNION; + wind_thresh = [ NA ]; + wind_logic = UNION; + + field = [ + { + name = "IWV"; + level = [ "L0" ]; + cat_thresh = [ >25.0 ]; + } + ]; + } message_type ^^^^^^^^^^^^ @@ -1592,54 +1592,54 @@ than one "message_type" entry is desired within the config file. For example: .. code-block:: none - fcst = { - censor_thresh = []; - censor_val = []; - cnt_thresh = [ NA ]; - cnt_logic = UNION; - wind_thresh = [ NA ]; - wind_logic = UNION; - - field = [ - { - message_type = [ "ADPUPA" ]; - sid_inc = []; - sid_exc = []; - name = "TMP"; - level = [ "P250", "P500", "P700", "P850", "P1000" ]; - cat_thresh = [ <=273.0 ]; - }, - { - message_type = [ "ADPSFC" ]; - sid_inc = []; - sid_exc = [ "KDEN", "KDET" ]; - name = "TMP"; - level = [ "Z2" ]; - cat_thresh = [ <=273.0 ]; - } - ]; - } + fcst = { + censor_thresh = []; + censor_val = []; + cnt_thresh = [ NA ]; + cnt_logic = UNION; + wind_thresh = [ NA ]; + wind_logic = UNION; + + field = [ + { + message_type = [ "ADPUPA" ]; + sid_inc = []; + sid_exc = []; + name = "TMP"; + level = [ "P250", "P500", "P700", "P850", "P1000" ]; + cat_thresh = [ <=273.0 ]; + }, + { + message_type = [ "ADPSFC" ]; + sid_inc = []; + sid_exc = [ "KDEN", "KDET" ]; + name = "TMP"; + level = [ "Z2" ]; + cat_thresh = [ <=273.0 ]; + } + ]; + } sid_inc and sid_exc ^^^^^^^^^^^^^^^^^^^ The "sid_inc" entry is an array of station ID groups indicating which -station ID's should be included in the verification task. If specified, -only those station ID's appearing in the list will be included. Note +station IDs should be included in the verification task. If specified, +only those station IDs appearing in the list will be included. Note that filtering by station ID may also be accomplished using the "mask.sid" option. However, when using the "sid_inc" option, statistics are reported separately for each masking region. The "sid_exc" entry is an array of station ID groups indicating which -station ID's should be excluded from the verification task. +station IDs should be excluded from the verification task. Each element in the "sid_inc" and "sid_exc" arrays is either the name of a single station ID or the full path to a station ID group file name. A station ID group file consists of a name for the group followed by a -list of station ID's. All of the station ID's indicated will be concatenated -into one long list of station ID's to be included or excluded. +list of station IDs. All of the station IDs indicated will be concatenated +into one long list of station IDs to be included or excluded. As with "message_type" above, the "sid_inc" and "sid_exc" settings can be -placed in the in the "field" array element to control which station ID's +placed in the "field" array element to control which station IDs are included or excluded for each verification task. .. code-block:: none @@ -1665,7 +1665,7 @@ field ^^^^^ The "field" entry is an array of dictionaries, specified the same way as those in the "fcst" and "obs" dictionaries. If the array has -length zero, not climatology data will be read and all climatology +length zero, no climatology data will be read and all climatology statistics will be written as missing data. Otherwise, the array length must match the length of "field" in the "fcst" and "obs" dictionaries. @@ -1680,9 +1680,9 @@ time_interp_method The "time_interp_method" entry specifies how the climatology data should be interpolated in time to the forecast valid time: - * NEAREST for data closest in time - * UW_MEAN for average of data before and after - * DW_MEAN for linear interpolation in time of data before and after + * NEAREST for data closest in time + * UW_MEAN for average of data before and after + * DW_MEAN for linear interpolation in time of data before and after day_interval ^^^^^^^^^^^^ @@ -1733,15 +1733,15 @@ configuration file context to use the same data for both. The "climo_mean" and assuming normality. These climatological distributions are used in two ways: (1) - To define climatological distribution percentiles thresholds (FCDP and - OCDP) which can be used as categorical (cat_thresh), continuous (cnt_thresh), - or wind speed (wind_thresh) thresholds. + To define climatological distribution percentile thresholds (FCDP and + OCDP) which can be used as categorical (cat_thresh), continuous (cnt_thresh), + or wind speed (wind_thresh) thresholds. (2) - To subset matched pairs into climatological bins based on where the - observation value falls within the observation climatological distribution. - See the "climo_cdf" dictionary. Note that only the observation climatology - data is used for this purpose, not the forecast climatology data. + To subset matched pairs into climatological bins based on where the + observation value falls within the observation climatological distribution. + See the "climo_cdf" dictionary. Note that only the observation climatology + data is used for this purpose, not the forecast climatology data. This dictionary is identical to the "climo_mean" dictionary described above but points to files containing climatological standard deviation values @@ -1779,7 +1779,7 @@ dictionaries, as shown below. climo_cdf --------- -The "climo_cdf" dictionary specifies how the the observation climatological +The "climo_cdf" dictionary specifies how the observation climatological mean ("climo_mean") and standard deviation ("climo_stdev") data are used to evaluate model performance relative to where the observation value falls within the observation climatological distribution. It can be set inside the @@ -1787,14 +1787,14 @@ within the observation climatological distribution. It can be set inside the dictionary consists of the following entries: (1) - The "cdf_bins" entry defines the climatological bins either as an integer - or an array of floats between 0 and 1. + The "cdf_bins" entry defines the climatological bins either as an integer + or an array of floats between 0 and 1. (2) - The "center_bins" entry may be set to TRUE or FALSE. + The "center_bins" entry may be set to TRUE or FALSE. (3) - The "write_bins" entry may be set to TRUE or FALSE. + The "write_bins" entry may be set to TRUE or FALSE. (4) The "direct_prob" entry may be set to TRUE or FALSE. @@ -1822,7 +1822,7 @@ an even number of bins can only be uncentered. For example: 4 uncentered bins (cdf_bins = 4; center_bins = FALSE;) yields: 0.0, 0.25, 0.50, 0.75, 1.0 5 uncentered bins (cdf_bins = 5; center_bins = FALSE;) yields: - 0.0, 0.2, 0.4, 0.6, 0.8, 0.9, 1.0 + 0.0, 0.2, 0.4, 0.6, 0.8, 1.0 5 centered bins (cdf_bins = 5; center_bins = TRUE;) yields: 0.0, 0.125, 0.375, 0.625, 0.875, 1.0 @@ -1867,7 +1867,7 @@ true, the climatological probability is computed directly from the climatological distribution at each point as the area to the left of the event threshold value. For greater-than or greater-than-or-equal-to thresholds, 1.0 minus the area is used. When "direct_prob" is false, the -"cdf_bins" values are sampled from climatological distribution. The probability +"cdf_bins" values are sampled from the climatological distribution. The probability is computed as the proportion of those samples which meet the threshold criteria. In this way, the number of bins impacts the resolution of the climatological probabilities. These derived probability values are used to compute the @@ -1909,13 +1909,13 @@ mask_missing_flag The "mask_missing_flag" entry specifies how missing data should be handled in the Wavelet-Stat and MODE tools: - * NONE to perform no masking of missing data + * NONE to perform no masking of missing data - * FCST to mask the forecast field with missing observation data + * FCST to mask the forecast field with missing observation data - * OBS to mask the observation field with missing forecast data + * OBS to mask the observation field with missing forecast data - * BOTH to mask both fields with missing data from the other + * BOTH to mask both fields with missing data from the other .. code-block:: none @@ -1929,7 +1929,7 @@ The "obs_window" entry is a dictionary specifying a beginning ("beg" entry) and ending ("end" entry) time offset values in seconds. It defines the time window over which observations are retained for scoring. These time offsets are defined relative to a reference time t, as [t+beg, t+end]. -In PB2NC, the reference time is the PREPBUFR files center time. In +In PB2NC, the reference time is the PREPBUFR file's center time. In Point-Stat and Ensemble-Stat, the reference time is the forecast valid time. .. code-block:: none @@ -1951,10 +1951,10 @@ used in the computation of statistics. .. note:: - Masking regions can be defined in a variety of ways, described below. - However, if no geographic masking regions are specified, the MET tools - automatically set "grid" equal to "FULL" to verify all data in the - entire input domain. + Masking regions can be defined in a variety of ways, described below. + However, if no geographic masking regions are specified, the MET tools + automatically set "grid" equal to "FULL" to verify all data in the + entire input domain. Masking regions may be specified in the following ways: @@ -1968,7 +1968,7 @@ three digit grid number. Supplying a value of "FULL" indicates that the verification should be performed over the entire grid on which the data resides. See: `ON388 - TABLE B, GRID IDENTIFICATION (PDS Octet 7), MASTER LIST OF NCEP STORAGE GRIDS, GRIB Edition 1 (FM92) `_. -The "grid" entry can be the gridded data file defining grid. +The "grid" entry can be the gridded data file defining the grid. poly ^^^^ @@ -1988,19 +1988,19 @@ These three options are described below: If providing an ASCII file containing the lat/lon points defining the mask polygon, the file must contain a name for the region followed by the latitude (degrees north) and longitude (degrees east) for each vertex of the polygon. - The values are separated by whitespace (e.g. spaces or newlines), and the + The values are separated by whitespace (e.g., spaces or newlines), and the first and last polygon points are connected. The general form is "poly_name lat1 lon1 lat2 lon2... latn lonn". Here is an example of a rectangle consisting of 4 points: .. code-block:: none - :caption: ASCII Rectangle Polygon Mask + :caption: ASCII Rectangle Polygon Mask - RECTANGLE - 25 -120 - 55 -120 - 55 -70 - 25 -70 + RECTANGLE + 25 -120 + 55 -120 + 55 -70 + 25 -70 Several masking polygons used by NCEP are predefined in the installed *share/met/poly* directory. Creating a new polygon is as @@ -2010,12 +2010,12 @@ These three options are described below: lat/lon polygon points are converted into x/y values in the grid. The lat/lon values for the observation points are also converted into x/y grid coordinates. The computations performed to check whether the - observation point falls within the polygon defined is done in x/y + observation point falls within the polygon defined are done in x/y grid space. .. code-block:: none - mask = { poly = [ "share/met/poly/CONUS.poly" ]; } + mask = { poly = [ "share/met/poly/CONUS.poly" ]; } * Option 2 - Gen-Vx-Mask output: @@ -2024,7 +2024,7 @@ These three options are described below: .. code-block:: none - mask = { poly = [ "/path/to/gen_vx_mask_output.nc" ]; } + mask = { poly = [ "/path/to/gen_vx_mask_output.nc" ]; } * Option 3 - Any gridded data file: @@ -2038,20 +2038,20 @@ These three options are described below: .. code-block:: none - mask = { poly = [ "/path/to/sample.grib {name = \"TMP\"; level = \"Z2\";} >273" ]; } + mask = { poly = [ "/path/to/sample.grib {name = \"TMP\"; level = \"Z2\";} >273" ]; } .. note:: The syntax for the Option 3 is complicated since it includes quotes embedded within another quoted string. Any such embedded quotes must - be escaped using a preceeding backslash character. + be escaped using a preceding backslash character. sid and llpnt ^^^^^^^^^^^^^ The "sid" entry is an array of strings which define groups of observation station ID's over which to compute statistics. Each station ID string can be followed by an -optional numeric weight enclosed in parenethesis and used by the "point_weight_flag" -configuration option. Each entry in the "sid" "array is either a filename or a +optional numeric weight enclosed in parentheses and used by the "point_weight_flag" +configuration option. Each entry in the "sid" array is either a filename or a comma-separated list. * For an ASCII filename, the strings contained within it are whitespace-separated. @@ -2059,7 +2059,7 @@ comma-separated list. ID's to be used. * For a comma-separated list, optionally use a colon to specify a name. For "MY_LIST:SID1(WGT1),SID2(WGT2)", name = MY_LIST which consists of - two station ID's (SID1 and SID2) and optional numeric weights (WGT1 and WGT2). + two station IDs (SID1 and SID2) and optional numeric weights (WGT1 and WGT2). * For a comma-separated list of length one with no name specified, the mask "name" and value are both set to the single station ID string. For "SID1", name = SID1 and value = SID1. @@ -2079,9 +2079,9 @@ longitude values meet this threshold criteria are used. A threshold set to "NA" always evaluates to true. The masking logic for processing point observations in Point-Stat and -Ensemble-Stat fall into two cateogries. The "sid" and "llpnt" options apply +Ensemble-Stat falls into two categories. The "sid" and "llpnt" options apply directly to the point observations. Only those observations for the specified -station id's are included in the "sid" masks. Only those observations meeting +station IDs are included in the "sid" masks. Only those observations meeting the latitude and longitude threshold criteria are included in the "llpnt" masks. @@ -2167,7 +2167,7 @@ Setting this variable to zero disables the computation of bootstrap confidence intervals, which may be necessary to run MET in realtime or near-realtime over large domains since bootstrapping is computationally expensive. Setting this variable to 1000 indicates that bootstrap -confidence interval should be computed over 1000 subsamples of the +confidence intervals should be computed over 1000 subsamples of the matched pairs. rng @@ -2187,7 +2187,7 @@ of bootstrap confidence intervals fully repeatable. When left empty the random number generator seed is chosen automatically which will lead to slightly different bootstrap confidence intervals being computed each time the data is run. Specifying a value here ensures that the bootstrap -confidence intervals will be reproducable over multiple runs on the same +confidence intervals will be reproducible over multiple runs on the same computing platform. .. code-block:: none @@ -2253,9 +2253,9 @@ For squares, a width of 2 defines a 2 x 2 box of grid points around the observation point (the 4 closest model grid points), while a width of 3 defines a 3 x 3 box of grid points around the observation point, and so on. For odd widths in grid-to-point comparisons -(i.e. Point-Stat), the interpolation area is centered on the model +(i.e., Point-Stat), the interpolation area is centered on the model grid point closest to the observation point. For grid-to-grid -comparisons (i.e. Grid-Stat), the width must be odd. +comparisons (i.e., Grid-Stat), the width must be odd. type.method """"""""""" @@ -2302,7 +2302,7 @@ applied to the points in the box: .. note:: Requesting the GEOG_MATCH interpolation method without providing any - land/sea mask (e.g. "land_mask.flag = FALSE") or topography data (e.g. + land/sea mask (e.g., "land_mask.flag = FALSE") or topography data (e.g., "topo_mask.flag = FALSE") results in a warning message. Without input geography data, GEOG_MATCH produces the same result as NEAREST. @@ -2341,7 +2341,7 @@ comma-separated lists of message types whose observations exist on land or water respectively. For point observations whose message type appears in the "LANDSF" entry, only interpolate using forecast grid points where land = TRUE. For point observations whose message type appears in the "WATERSF" entry, only interpolate -using forecast grids points where land = FALSE. By default, "ADPSFC" and "MSONET" +using forecast grid points where land = FALSE. By default, "ADPSFC" and "MSONET" message types exist over land while the "SFCSHP" message type exists over water. .. code-block:: none @@ -2353,7 +2353,6 @@ message types exist over land while the "SFCSHP" message type exists over water. ... ]; -The "topo_mask.flag", "topo_mask.use_obs_thresh", and "topo_mask.interp_fcst_thresh" The "land_mask.flag" entry may be set separately in each "obs.field" entry. .. code-block:: none @@ -2371,7 +2370,7 @@ topo_mask The "topo_mask" dictionary defines the model topography field used when verifying at the surface. The flag entry enables/disables this logic. -Only use point observations where the model topography minus station elevation +Only use point observations where the model topography minus station elevation difference meets the "use_obs_thresh" threshold entry. For the observations kept, when interpolating forecast data to the observation location, only use forecast grid points where the topo minus station @@ -2482,7 +2481,7 @@ forecast value using the model topography height and convert the observation value using the station elevation. The "thresh" option specifies the valid range of values to be converted. Values not meeting this threshold criteria are left unchanged. The default threshold of "NA" -always evaulates to true, but it can be set to avoid converting flag values. For example, +always evaluates to true, but it can be set to avoid converting flag values. For example, set "thresh = ne99999;" to avoid converting a cloud base height flag value of 99999 which may indicate clear sky. The "msl_to_agl" entry is a boolean. When "TRUE", the elevation correction is @@ -2535,7 +2534,7 @@ the ratio of the nearby forecast values that meet the threshold criteria. Point-Stat evaluates those fractional coverage values as if they were a probability forecast. When applying HiRA, users should enable the matched pair (MPR), probabilistic (PCT, PSTD, PJC, or PRC), or ensemble statistics -(ECNT or PRS) line types in the output_flag dictionary. The number of +(ECNT or RPS) line types in the output_flag dictionary. The number of probabilistic HiRA output lines is determined by the number of categorical forecast thresholds and HiRA neighborhood widths chosen. This dictionary may include the following entries: @@ -2666,13 +2665,13 @@ The "nc_pairs_flag" can be set either to a boolean value or a dictionary in either Grid-Stat, Wavelet-Stat or MODE. The dictionary (with slightly different entries for the various tools ... see the default config files) has individual boolean settings turning on or off the writing out of the -various fields in the netcdf output file for the tool. Setting all -dictionary entries to false means the netcdf file will not be generated. +various fields in the NetCDF output file for the tool. Setting all +dictionary entries to false means the NetCDF file will not be generated. "nc_pairs_flag" can also be set to a boolean value. In this case, a value of true means to just accept the default settings (which will turn on the output of all the different fields). A value of false means no -netcdf output will be generated. +NetCDF output will be generated. .. code-block:: none @@ -2686,7 +2685,7 @@ netcdf output will be generated. nbrhd = FALSE; fourier = FALSE; gradient = FALSE; - distance_map = FLASE; + distance_map = FALSE; apply_mask = TRUE; } @@ -2726,7 +2725,7 @@ For example: .. note:: - Prior to MET version 9.0.0, this option was named "nc_pairs_var_str",' + Prior to MET version 9.0.0, this option was named "nc_pairs_var_str", which is now deprecated. .. code-block:: none @@ -2763,7 +2762,7 @@ Three grid weighting options are currently supported: * NONE to disable grid weighting using a constant weight of 1.0 (default). * COS_LAT to define the weight as the cosine of the grid point latitude. - This an approximation for grid box area used by NCEP and WMO. + This is an approximation for grid box area used by NCEP and WMO. * AREA to define the weight as the true area of the grid box (km^2). @@ -2802,7 +2801,7 @@ It is not applied for grid-to-grid verification which is controlled by the "grid_weight_flag" option. It can only be defined once at the highest level of config file context and applies to all verification tasks for that run. -While only one point weighting option is currently supported, additional +The following point weighting options are currently supported, and additional methods are planned for future versions: * NONE to disable point weighting using a constant weight of 1.0 (default). @@ -2825,7 +2824,7 @@ divided by the CTC or MCTC table dimension. For example, for a 2x2 CTC table, the default hss_ec_value is 1.0 / 2 = 0.5. For a 4x4 MCTC table, the default hss_ec_value is 1.0 / 4 = 0.25. -If set, it must greater than or equal to 0.0 and less than 1.0. A value of +If set, it must be greater than or equal to 0.0 and less than 1.0. A value of 0.0 produces an HSS_EC statistic equal to the Accuracy statistic. .. code-block:: none @@ -3033,11 +3032,11 @@ For example: .. code-block:: none - beg = "00"; - end = "235959"; - step = 300; - width = 600; - width = { beg = -300; end = 300; } + beg = "00"; + end = "235959"; + step = 300; + width = 600; + width = { beg = -300; end = 300; } This example does a 10-minute time summary every 5 minutes throughout the day. The first interval will be from 23:55:00 the previous day through @@ -3069,7 +3068,7 @@ The "vld_freq" and "vld_thresh" options may be used to require that a certain ratio of observations must be present and contain valid data within the time window in order for a summary value to be computed. The "vld_freq" entry defines the expected observation frequency in seconds. For example, when -summarizing 1-minute data (vld_freq = 60) over a 30 minute time window, +summarizing 1-minute data (vld_freq = 60) over a 30-minute time window, setting "vld_thresh = 0.5" requires that at least 15 of the 30 expected observations be present and valid for a summary value to be written. The default "vld_thresh = 0.0" setting will skip over this logic. @@ -3214,13 +3213,13 @@ centered on the current point, and the width array specifies the candidate neighborhood sizes to be considered. Each width specifies the width of the square or diameter of the circle as an odd integer. The vld_thresh entry is a number between 0 and 1 specifying the required ratio of valid data in the -neighborhood for an output value to be computed. The alpha entry is number +neighborhood for an output value to be computed. The alpha entry is a number between 0 and 1 specifying the EAS distance criteria. For each grid point, the smallest width for which the distance criteria is satisfied is used. If the distance criteria is never satisfied, the largest width is used. The gaussian_dx and gaussian_radius entries define the Gaussian smoother -which is applied to be raw EAS probabilities. +which is applied to the raw EAS probabilities. If ensemble_flag.eas is set to TRUE, EAS probabilities are written for each categorical threshold (cat_thresh) specified. If ensemble_flag.eas_width is @@ -3240,7 +3239,7 @@ set to TRUE, the widths chosen to compute EAS are written for each threshold. ensemble_flag ^^^^^^^^^^^^^ -The "ensemble_flag" entry is a dictionary of boolean value indicating +The "ensemble_flag" entry is a dictionary of boolean values indicating which ensemble products should be generated: * "latlon" for a grid of the Latitude and Longitude fields @@ -3354,7 +3353,7 @@ empty string, meaning that no customization is applied to the output variable names. When the Ensemble-Stat config file contains two fields with the same name and level value, this entry is used to make the resulting variable names unique. -e.g. nc_var_str = "MIN"; +e.g., nc_var_str = "MIN"; .. code-block:: none @@ -3474,7 +3473,7 @@ lines are read in and processed. The MODE line options are numerous. They fall into seven categories: toggles, multiple set string options, multiple set integer options, integer max/min options, date/time max/min options, floating-point max/min options, and miscellaneous options. **In order to be -applied, the options must be uncommented (i.e. remove the "//" marks) before +applied, the options must be uncommented (i.e., remove the "//" marks) before running.** These options are described in subsequent sections. Please note that this configuration file is processed differently than the other config files. @@ -3486,7 +3485,7 @@ Toggles The MODE line options described in this section are shown in pairs. These toggles represent parameters that can have only one (or none) of two values. Any of these toggles may be left unspecified. However, if neither -option for toggle is indicated, the analysis will produce results that +option for a toggle is indicated, the analysis will produce results that combine data from both toggles. This may produce unintended results. @@ -3531,13 +3530,13 @@ separated by spaces. Each of these options must be indicated as a string. String values that include spaces may be used by enclosing the string in quotation marks. -This options specifies which model to use +This option specifies which model to use. .. code-block:: none // model = []; -These two options specify thresholds for forecast and observations objects to +These two options specify thresholds for forecast and observation objects to be used in the analysis, respectively. .. code-block:: none @@ -3597,7 +3596,7 @@ time. // fcst_accum = []; // obs_accum = []; -These options indicate the convolution radius used for forecast of observed +These options indicate the convolution radius used for forecast or observed objects, respectively. .. code-block:: none @@ -3656,11 +3655,11 @@ Date/time max/min Options ^^^^^^^^^^^^^^^^^^^^^^^^^ These options set limits on various date/time attributes. The values can be specified in one of three ways: First, the -options may be indicated by a string of the form YYYMMDD_HHMMSS. This +options may be indicated by a string of the form YYYYMMDD_HHMMSS. This specifies a complete calendar date and time. Second, they may be indicated -by a string of the form YYYYMMMDD_HH. Here, the minutes and seconds are +by a string of the form YYYYMMDD_HH. Here, the minutes and seconds are assumed to be zero. The third way of indicating date/time attributes is by a -string of the form YYYMMDD. Here, hours, minutes, and seconds are assumed to +string of the form YYYYMMDD. Here, hours, minutes, and seconds are assumed to be zero. These options indicate minimum/maximum values for the forecast valid time. @@ -3847,7 +3846,7 @@ fcst/obs.filter_attr_name and fcst/obs.filter_attr_thresh """"""""""""""""""""""""""""""""""""""""""""""""""""""""" The "filter_attr_name" and "filter_attr_thresh" entries are arrays of the same length which specify object filtering criteria. By default, no -object filtering criteria is defined. +object filtering criteria are defined. The "filter_attr_name" entry is an array of strings specifying the MODE output header column names for the object attributes of interest, such @@ -3860,7 +3859,7 @@ The "filter_attr_thresh" entry is an array of thresholds for the object attributes. Any simple objects not meeting all of these filtering criteria are discarded. -Note that the "area_thresh" and "inten_perc_thresh" entries form +Note that the "area_thresh" and "inten_perc_thresh" entries from earlier versions of MODE are replaced by these options and are now deprecated. @@ -3879,14 +3878,14 @@ fcst/obs.merge_flag """"""""""""""""""" The "merge_flag" entry specifies the merging methods to be applied: - * NONE for no merging + * NONE for no merging - * THRESH for the double-threshold merging method. Merge objects - that would be part of the same object at the lower threshold. + * THRESH for the double-threshold merging method. Merge objects + that would be part of the same object at the lower threshold. - * ENGINE for the fuzzy logic approach comparing the field to itself + * ENGINE for the fuzzy logic approach comparing the field to itself - * BOTH for both the double-threshold and engine merging methods + * BOTH for both the double-threshold and engine merging methods .. code-block:: none @@ -3913,7 +3912,7 @@ grid_res The "grid_res" entry is the nominal spacing for each grid square in kilometers. The variable is not used directly in the code, but subsequent variables in the configuration files are defined in terms of it. Therefore, -setting the appropriately will help ensure that appropriate default values +setting it appropriately will help ensure that appropriate default values are used for these variables. .. code-block:: none @@ -4113,7 +4112,7 @@ following criteria: (1) by message type: supply a list of PREPBUFR message types to retain -(2) by station id: supply a list of observation stations to retain +(2) by station ID: supply a list of observation stations to retain (3) by valid time: supply the beginning and ending time offset values in the obs_window entry described above. @@ -4127,7 +4126,7 @@ following criteria: (6) by report type: supply a list of report types to retain using pb_report_type and in_report_type entries described below -(7) by instrument type: supply a list of instrument type to +(7) by instrument type: supply a list of instrument types to retain (8) by vertical level: supply beg/end vertical levels using the @@ -4190,8 +4189,8 @@ For example: station_id ^^^^^^^^^^ -The "station_id" entry is an array of station ids to be retained or -the filename which contains station ids. An array of station ids +The "station_id" entry is an array of station IDs to be retained or +the filename which contains station IDs. An array of station IDs contains a comma-separated list. An empty list indicates that all stations should be retained. @@ -4205,7 +4204,7 @@ elevation_range ^^^^^^^^^^^^^^^ The "elevation_range" entry is a dictionary which contains "beg" and "end" -entries specifying the range of observing locations elevations to be +entries specifying the range of observing location elevations to be retained. .. code-block:: none @@ -4336,7 +4335,7 @@ command line option to see the list of available observation variables. obs_bufr_map ^^^^^^^^^^^^ -Mapping of input BUFR variable names to output variables names. +Mapping of input BUFR variable names to output variable names. The default PREPBUFR map, obs_prepbufr_map, is appended to this map. Users may choose to rename BUFR variables to match the naming convention of the forecast the observation is used to verify. @@ -4386,11 +4385,11 @@ See `Code table for observation quality markers `_". +The MET tools use the following attributes and variables for input "`CF Compliant NetCDF data `_". 1. The global attribute "Conventions". @@ -63,7 +63,7 @@ The MET tools use following attributes and variables for input "`CF Compliant Ne 7. (Optional) the "`vertical coordinate `_" variable must include a supported units attribute to allow range selection. -MET processes the CF-Compliant gridded NetCDF files with the projection information. The CF-Compliant NetCDF is defined by the global attribute "Conventions" whose value begins with "CF-" ("CF-"). The global attribute "Conventions" is mandatory. MET accepts the variation of this attribute ("conventions" and "CONVENTIONS"). The value should be started with "CF-" and followed by the version number. MET accepts the attribute value that begins with "CF " ("CF" and a space instead of a hyphen) or "COARDS". +MET processes the CF-Compliant gridded NetCDF files with the projection information. The CF-Compliant NetCDF is defined by the global attribute "Conventions" whose value begins with "CF-" ("CF-"). The global attribute "Conventions" is mandatory. MET accepts the variation of this attribute ("conventions" and "CONVENTIONS"). The value should start with "CF-" and followed by the version number. MET accepts the attribute value that begins with "CF " ("CF" and a space instead of a hyphen) or "COARDS". The grid mapping variable contains the projection information. The grid mapping variable can be found by looking at the variable attribute "grid_mapping" from the data variables. The "standard_name" attribute is used to filter out the coordinate variables like time, latitude, and longitude variables. The value of the "grid_mapping" attribute is the name of the grid mapping variable. Four projections are supported with grid mapping variables: latitude_longitude, lambert_conformal_conic, polar_stereographic, and geostationary. In case of the latitude_longitude projection, the latitude and longitude variable names should be the same as the dimension names and the "units" attribute should be valid. @@ -73,37 +73,37 @@ Here are examples for the grid mapping variable ("edr" is the data variable): .. code-block:: none - float edr(time, z, lat, lon) ; - edr:units = "m^(2/3) s^-1" ; - edr:long_name = "Median eddy dissipation rate" ; - edr:coordinates = "lat lon" ; - edr:_FillValue = -9999.f ; - edr:grid_mapping = "grid_mapping" ; - int grid_mapping ; - grid_mapping:grid_mapping_name = "latitude_longitude" ; - grid_mapping:semi_major_axis = 6371000. ; - grid_mapping:inverse_flattening = 0 ; + float edr(time, z, lat, lon) ; + edr:units = "m^(2/3) s^-1" ; + edr:long_name = "Median eddy dissipation rate" ; + edr:coordinates = "lat lon" ; + edr:_FillValue = -9999.f ; + edr:grid_mapping = "grid_mapping" ; + int grid_mapping ; + grid_mapping:grid_mapping_name = "latitude_longitude" ; + grid_mapping:semi_major_axis = 6371000. ; + grid_mapping:inverse_flattening = 0 ; **Example 2: grid mapping for lambert_conformal_conic projection** .. code-block:: none - float edr(time, z, y, x) ; - edr:units = "m^(2/3) s^-1" ; - edr:long_name = "Eddy dissipation rate" ; - edr:coordinates = "lat lon" ; - edr:_FillValue = -9999.f ; - edr:grid_mapping = "grid_mapping" ; - int grid_mapping ; - grid_mapping:grid_mapping_name = "lambert_conformal_conic" ; - grid_mapping:standard_parallel = 25. ; - grid_mapping:longitude_of_central_meridian = -95. ; - grid_mapping:latitude_of_projection_origin = 25. ; - grid_mapping:false_easting = 0 ; - grid_mapping:false_northing = 0 ; - grid_mapping:GRIB_earth_shape = "spherical" ; - grid_mapping:GRIB_earth_shape_code = 0 ; + float edr(time, z, y, x) ; + edr:units = "m^(2/3) s^-1" ; + edr:long_name = "Eddy dissipation rate" ; + edr:coordinates = "lat lon" ; + edr:_FillValue = -9999.f ; + edr:grid_mapping = "grid_mapping" ; + int grid_mapping ; + grid_mapping:grid_mapping_name = "lambert_conformal_conic" ; + grid_mapping:standard_parallel = 25. ; + grid_mapping:longitude_of_central_meridian = -95. ; + grid_mapping:latitude_of_projection_origin = 25. ; + grid_mapping:false_easting = 0 ; + grid_mapping:false_northing = 0 ; + grid_mapping:GRIB_earth_shape = "spherical" ; + grid_mapping:GRIB_earth_shape_code = 0 ; When the grid mapping variable is not available, MET can detect either a latitude_longitude or rotated_latitude_longitude projection. It detects the latitude_longitude projection in the following order: @@ -113,7 +113,7 @@ When the grid mapping variable is not available, MET can detect either a latitud 3. the lat/lon projection from the latitude and longitude variables by the "standard_name" attribute -MET is looking for variables with the same name as the dimension and checking the "units" attribute to find the latitude and longitude variables. The valid "units" strings are listed in the table below. MET accepts the variable "tlat" and "tlon" if the dimension names are "nlat" and "nlon”. +MET is looking for variables with the same name as the dimension and checking the "units" attribute to find the latitude and longitude variables. The valid "units" strings are listed in the table below. MET accepts the variable "tlat" and "tlon" if the dimension names are "nlat" and "nlon". If there are no latitude and longitude variables from dimensions, MET gets coordinate variable names from the "coordinates" attribute. The matching coordinate variables should have the proper "units" attribute. @@ -133,7 +133,7 @@ For rotated_latitude_longitude projections, MET detects the projection using the 3. Checking to see if the standard name attribute is called grid_latitude for latitude variables and grid_longitude for the longitude variable. -The latitude and longitude variables must be one dimensional and with their size matching the corresponding dimension for latitude_longitude and rotated_latitude_longitude grids. +The latitude and longitude variables must be one-dimensional and with their size matching the corresponding dimension for latitude_longitude and rotated_latitude_longitude grids. .. list-table:: Valid strings for the "units" attribute. :widths: auto @@ -168,7 +168,7 @@ Performance with NetCDF Input Data There is no limitation on the NetCDF file size. The size of the data variables matters more than the file size. The NetCDF API loads the metadata first upon opening the NetCDF file. It's similar for accessing data variables. There are two API calls: getting the metadata and getting the actual data. The memory is allocated and consumed at the second API call (getting the actual data). -The dimensions of the data variables matter. MET requests the NetCDF data needs based on: 1) loading and processing a data plane, and 2) loading and processing the next data plane. This means an extra step for slicing with one more dimension in the NetCDF input data. The performance is quite different if the compression is enabled with high resolution data. NetCDF does compression per variable. The variables can have different compression levels (0 to 9). A value of 0 means no compression, and 9 is the highest level of compression possible. The number for decompression is the same between one more and one less dimension NetCDF input files (combined VS separated). The difference is the amount of data to be decompressed which requires more memory. For example, let's assume the time dimension is 30. NetCDF data with one less dimension (no time dimension) does decompression 30 times for nx by ny dataset. NetCDF with one more dimension does compression 30 times for 30 by nx by ny dataset and slicing for target time offset. So it's better to have multiple NetCDF files with one less dimension than a big file with bigger variable data if compressed. If the compression is not enabled, the file size will be much bigger requiring more disk space. +The dimensions of the data variables matter. MET reads NetCDF data one data plane at a time: it loads and processes one data plane, and then loads and processes the next one. When the input variable has an extra dimension, this requires an extra slicing step. The performance is quite different if the compression is enabled with high resolution data. NetCDF does compression per variable. The variables can have different compression levels (0 to 9). A value of 0 means no compression, and 9 is the highest level of compression possible. The number of decompressions is the same for NetCDF input files with one more or one less dimension (combined vs. separated). The difference is the amount of data to be decompressed which requires more memory. For example, let's assume the time dimension is 30. NetCDF data with one less dimension (no time dimension) does decompression 30 times for an nx by ny dataset. NetCDF data with one more dimension does decompression 30 times for a 30 by nx by ny dataset, plus slicing for the target time offset. So it's better to have multiple NetCDF files with one less dimension than a big file with bigger variable data if compressed. If the compression is not enabled, the file size will be much bigger requiring more disk space. .. _Intermediate data formats: @@ -188,13 +188,13 @@ The STAT format consists of tabular ASCII data that can be easily read by many a For this reason, ASCII output is also available as an alternative for these tools. The ASCII files contain exactly the same output as the STAT files but each STAT line type is grouped into a single ASCII file with a column header row making the output more human-readable. The configuration files control which line types are output and whether or not the optional ASCII files are generated. -The MODE tool creates two ASCII output files as well (although they are not in a STAT format). It generates an ASCII file containing contingency table counts and statistics comparing the model and observation fields being compared. The MODE tool also generates a second ASCII file containing all of the attributes for the single objects and pairs of objects. Each line in this file contains the same number of columns, and those columns not applicable to a given line type contain fill data. Similarly, the MTD tool writes one ASCII output file for 2D objects attributes and four ASCII output files for 3D object attributes. +The MODE tool creates two ASCII output files as well (although they are not in a STAT format). It generates an ASCII file containing contingency table counts and statistics comparing the model and observation fields being compared. The MODE tool also generates a second ASCII file containing all of the attributes for the single objects and pairs of objects. Each line in this file contains the same number of columns, and those columns not applicable to a given line type contain fill data. Similarly, the MTD tool writes one ASCII output file for 2D object attributes and four ASCII output files for 3D object attributes. The TC-Pairs and TC-Stat utilities produce ASCII output, similar in style to the STAT files, but with TC relevant fields. -Many of the tools generate gridded NetCDF output. Generally, this output acts as input to other MET tools or plotting programs. The point observation preprocessing tools produce NetCDF output as input to the statistics tools. Full details of the contents of the NetCDF files is found in :numref:`Data format summary` below. +Many of the tools generate gridded NetCDF output. Generally, this output acts as input to other MET tools or plotting programs. The point observation preprocessing tools produce NetCDF output as input to the statistics tools. Full details of the contents of the NetCDF files are found in :numref:`Data format summary` below. -The MODE, Wavelet-Stat and plotting tools produce PostScript plots summarizing the spatial approach used in the verification. The PostScript plots are generated using internal libraries and do not depend on an external plotting package. The MODE plots contain several summary pages at the beginning, but the total number of pages will depend on the merging options chosen. Additional pages will be created if merging is performed using the double thresholding or fuzzy engine merging techniques for the forecast and observation fields. The number of pages in the Wavelet-Stat plots depend on the number of masking tiles used and the dimension of those tiles. The first summary page is followed by plots for the wavelet decomposition of the forecast and observation fields. The generation of these PostScript output files can be disabled using command line options. +The MODE, Wavelet-Stat and plotting tools produce PostScript plots summarizing the spatial approach used in the verification. The PostScript plots are generated using internal libraries and do not depend on an external plotting package. The MODE plots contain several summary pages at the beginning, but the total number of pages will depend on the merging options chosen. Additional pages will be created if merging is performed using the double thresholding or fuzzy engine merging techniques for the forecast and observation fields. The number of pages in the Wavelet-Stat plots depends on the number of masking tiles used and the dimension of those tiles. The first summary page is followed by plots for the wavelet decomposition of the forecast and observation fields. The generation of these PostScript output files can be disabled using command line options. Users can use the optional plotting utilities Plot-Data-Plane and Plot-Point-Obs to produce graphics showing forecast and observation data. @@ -284,7 +284,7 @@ The following is a summary of the input and output formats for each of the tools * **Output**: One STAT file containing all of the requested line types, several ASCII files for each line type requested, and one NetCDF file containing the matched pair data and difference field for each verification region and variable type/level being verified. -#. **Ensemble Stat Tool** +#. **Ensemble-Stat Tool** * **Input**: An arbitrary number of gridded model files, one or more gridded and/or point observation files, and one configuration file. Point and gridded observations are both accepted. @@ -310,7 +310,7 @@ The following is a summary of the input and output formats for each of the tools #. **Stat-Analysis Tool** - * **Input**: One or more STAT files output from the Point-Stat, Grid-Stat, Ensemble Stat, Wavelet-Stat, or TC-Gen tools and, optionally, one configuration file containing specifications for the analysis job(s) to be run on the STAT data. + * **Input**: One or more STAT files output from the Point-Stat, Grid-Stat, Ensemble-Stat, Wavelet-Stat, or TC-Gen tools and, optionally, one configuration file containing specifications for the analysis job(s) to be run on the STAT data. * **Output**: ASCII output of the analysis jobs is printed to the screen unless redirected to a file using the "-out" option or redirected to a STAT output file using the "-out_stat" option. @@ -382,7 +382,7 @@ The following is a summary of the input and output formats for each of the tools #. **Plot-Point-Obs Tool** - * **Input**: One NetCDF file containing point observation from the ASCII2NC, PB2NC, MADIS2NC, or LIDAR2NC tool. + * **Input**: One NetCDF file containing point observations from the ASCII2NC, PB2NC, MADIS2NC, or LIDAR2NC tool. * **Output**: One postscript file containing a plot of the requested field. @@ -403,7 +403,7 @@ The following is a summary of the input and output formats for each of the tools Output Column Types =================== -MET ASCII output files are in a whitespace-separated columnar format. The contents of each column are described for the tools which generate that output type, including section :numref:`point_stat-output` for the Point-Stat tool. Each individual column has an associated data type (i.e. "VERSION" is a string, as seen in table :numref:`table_PS_header_info_point-stat_out`). This data type information is useful when reading MET output using Python. The data types of the columns for each MET line type can be found in *data/table_files/met_column_types.json*. +MET ASCII output files are in a whitespace-separated columnar format. The contents of each column are described for the tools which generate that output type, including section :numref:`point_stat-output` for the Point-Stat tool. Each individual column has an associated data type (i.e., "VERSION" is a string, as seen in table :numref:`table_PS_header_info_point-stat_out`). This data type information is useful when reading MET output using Python. The data types of the columns for each MET line type can be found in *data/table_files/met_column_types.json*. .. _Configuration File Details: diff --git a/docs/Users_Guide/ensemble-stat.rst b/docs/Users_Guide/ensemble-stat.rst index 61d53777e9..6659fedffd 100644 --- a/docs/Users_Guide/ensemble-stat.rst +++ b/docs/Users_Guide/ensemble-stat.rst @@ -48,9 +48,9 @@ Often, the goal of ensemble forecasting is to reproduce the distribution of obse The relative position (RELP) is a count of the number of times each ensemble member is closest to the observation. For stochastic or randomly derived ensembles, this statistic is meaningless. For specified ensemble members, however, it can assist users in determining if any ensemble member is performing consistently better or worse than the others. -The ranked probability score (RPS) is included in the Ranked Probability Score (RPS) line type. It is the mean of the Brier scores computed from ensemble probabilities derived for each probability category threshold (prob_cat_thresh) specified in the configuration file. The continuous ranked probability score (CRPS) is the average the distance between the forecast (ensemble) cumulative distribution function and the observation cumulative distribution function. It is an analog of the Brier score, but for continuous forecast and observation fields. The CRPS statistic is computed using two methods: assuming a normal distribution defined by the ensemble mean and spread (:ref:`Gneiting et al., 2004 `) and using the empirical ensemble distribution (:ref:`Hersbach, 2000 `). The CRPS statistic using the empirical ensemble distribution can be adjusted (bias corrected) by subtracting 1/(2*m) times the mean absolute difference of the ensemble members, where m is the ensemble size. This is reported as a separate statistic called CRPS_EMP_FAIR. The empirical CRPS and its fair version are included in the Ensemble Continuous Statistics (ECNT) line type, along with other statistics quantifying the ensemble spread and ensemble mean skill. +The ranked probability score (RPS) is included in the Ranked Probability Score (RPS) line type. It is the mean of the Brier scores computed from ensemble probabilities derived for each probability category threshold (prob_cat_thresh) specified in the configuration file. The continuous ranked probability score (CRPS) is the average distance between the forecast (ensemble) cumulative distribution function and the observation cumulative distribution function. It is an analog of the Brier score, but for continuous forecast and observation fields. The CRPS statistic is computed using two methods: assuming a normal distribution defined by the ensemble mean and spread (:ref:`Gneiting et al., 2004 `) and using the empirical ensemble distribution (:ref:`Hersbach, 2000 `). The CRPS statistic using the empirical ensemble distribution can be adjusted (bias corrected) by subtracting 1/(2*m) times the mean absolute difference of the ensemble members, where m is the ensemble size. This is reported as a separate statistic called CRPS_EMP_FAIR. The empirical CRPS and its fair version are included in the Ensemble Continuous Statistics (ECNT) line type, along with other statistics quantifying the ensemble spread and ensemble mean skill. -The Ensemble-Stat tool can derive ensemble relative frequencies and verify them as probability forecasts all in the same run. Note however that these simple ensemble relative frequencies are not actually calibrated probability forecasts. If probabilistic line types are requested (output_flag), this logic is applied to each pair of fields listed in the forecast (fcst) and observation (obs) dictionaries of the configuration file. Each probability category threshold (prob_cat_thresh) listed for the forecast field is applied to the input ensemble members to derive a relative frequency forecast. The probability category threshold (prob_cat_thresh) parsed from the corresponding observation entry is applied to the (gridded or point) observations to determine whether or not the event actually occurred. The paired ensemble relative frequencies and observation events are used to populate an Nx2 probabilistic contingency table. The dimension of that table is determined by the probability PCT threshold (prob_pct_thresh) configuration file option parsed from the forecast dictionary. All probabilistic output types requested are derived from this Nx2 table and written to the ascii output files. Note that the FCST_VAR name header column is automatically reset as "PROB({FCST_VAR}{THRESH})" where {FCST_VAR} is the current field being evaluated and {THRESH} is the threshold that was applied. +The Ensemble-Stat tool can derive ensemble relative frequencies and verify them as probability forecasts all in the same run. Note however that these simple ensemble relative frequencies are not actually calibrated probability forecasts. If probabilistic line types are requested (output_flag), this logic is applied to each pair of fields listed in the forecast (fcst) and observation (obs) dictionaries of the configuration file. Each probability category threshold (prob_cat_thresh) listed for the forecast field is applied to the input ensemble members to derive a relative frequency forecast. The probability category threshold (prob_cat_thresh) parsed from the corresponding observation entry is applied to the (gridded or point) observations to determine whether or not the event actually occurred. The paired ensemble relative frequencies and observation events are used to populate an Nx2 probabilistic contingency table. The dimension of that table is determined by the probability PCT threshold (prob_pct_thresh) configuration file option parsed from the forecast dictionary. All probabilistic output types requested are derived from this Nx2 table and written to the ASCII output files. Note that the FCST_VAR name header column is automatically reset as "PROB({FCST_VAR}{THRESH})" where {FCST_VAR} is the current field being evaluated and {THRESH} is the threshold that was applied. Note that if no probability category thresholds (prob_cat_thresh) are defined, but climatological mean and standard deviation data is provided along with climatological bins, climatological distribution percentile thresholds are automatically derived and used to compute probabilistic outputs. @@ -59,20 +59,20 @@ Climatology Data The Ensemble-Stat output includes at least three statistics computed relative to external climatology data. The climatology is defined by mean and standard deviation fields, and typically both are required in the computation of ensemble skill score statistics. MET assumes that the climatology follows a normal distribution, defined by the mean and standard deviation at each point. -When computing the CRPS skill score for (:ref:`Gneiting et al., 2004 `) the reference CRPS statistic is computed using the climatological mean and standard deviation directly. When computing the CRPS skill score for (:ref:`Hersbach, 2000 `) the reference CRPS statistic is computed by selecting equal-area-spaced values from the assumed normal climatological distribution. The number of points selected is determined by the *cdf_bins* setting in the *climo_cdf* dictionary. The reference CRPS is computed empirically from this ensemble of climatology values. If the number bins is set to 1, the climatological CRPS is computed using only the climatological mean value. In this way, the empirical CRPSS may be computed relative to a single model rather than a climatological distribution. +When computing the CRPS skill score for (:ref:`Gneiting et al., 2004 `) the reference CRPS statistic is computed using the climatological mean and standard deviation directly. When computing the CRPS skill score for (:ref:`Hersbach, 2000 `) the reference CRPS statistic is computed by selecting equal-area-spaced values from the assumed normal climatological distribution. The number of points selected is determined by the *cdf_bins* setting in the *climo_cdf* dictionary. The reference CRPS is computed empirically from this ensemble of climatology values. If the number of bins is set to 1, the climatological CRPS is computed using only the climatological mean value. In this way, the empirical CRPSS may be computed relative to a single model rather than a climatological distribution. -The climatological distribution is also used for the RPSS. The forecast RPS statistic is computed from a probabilistic contingency table in which the probabilities are derived from the ensemble member values. In a simliar fashion, the climatogical probability for each observed value is derived from the climatological distribution. The area of the distribution to the left of the observed value is interpreted as the climatological probability. These climatological probabilities are also evaluated using a probabilistic contingency table from which the reference RPS score is computed. The skill scores are derived by comparing the forecast statistic to the reference climatology statistic. +The climatological distribution is also used for the RPSS. The forecast RPS statistic is computed from a probabilistic contingency table in which the probabilities are derived from the ensemble member values. In a similar fashion, the climatological probability for each observed value is derived from the climatological distribution. The area of the distribution to the left of the observed value is interpreted as the climatological probability. These climatological probabilities are also evaluated using a probabilistic contingency table from which the reference RPS score is computed. The skill scores are derived by comparing the forecast statistic to the reference climatology statistic. -The Ensemble-Stat tool also allows the computation of RPS and RPSS utilizing an ensemble forecast in probabilistic space. This unique ability is for users with ensemble data that is formulated relative to climatology; it does not require the use of any climatological datasets, instead relying on the assumption that each ensemble member has a climatologically equal chance of occurring. Each ensemble member's field should contain values in the range [0, 1] or [0, 100]. However, when MET encounters a probability field with a range [0, 100], it will automatically rescale it to be [0, 1]. The sum of all ensemble member fields should equal 1 (if range is [0, 1]) or 100 (if range is [0, 100]). When calculating RPS, the ensemble member field values are cumulatively summed in the order that they are evaluated by Ensemble-Stat. Each of these sums is then used to calculate a squared probability error relative to observations and their appropriate thresolds. Note that it is expected each observation point or observation gridpoint will indicate exactly one climatological probability bin where the observation occurred. The cumulative sum and squared probability errors are also calculated for climatology using an even, constant probability bin width that is equal to 1 divided by the number of ensemble members (ex. three ensemble members are evaluated with three climatology probability bins of width 0.333). The accompanying skill score is derived by comparing the forecast RPS to the reference climatology RPS. +The Ensemble-Stat tool also allows the computation of RPS and RPSS utilizing an ensemble forecast in probabilistic space. This unique ability is for users with ensemble data that is formulated relative to climatology; it does not require the use of any climatological datasets, instead relying on the assumption that each ensemble member has a climatologically equal chance of occurring. Each ensemble member's field should contain values in the range [0, 1] or [0, 100]. However, when MET encounters a probability field with a range [0, 100], it will automatically rescale it to be [0, 1]. The sum of all ensemble member fields should equal 1 (if range is [0, 1]) or 100 (if range is [0, 100]). When calculating RPS, the ensemble member field values are cumulatively summed in the order that they are evaluated by Ensemble-Stat. Each of these sums is then used to calculate a squared probability error relative to observations and their appropriate thresholds. Note that it is expected each observation point or observation gridpoint will indicate exactly one climatological probability bin where the observation occurred. The cumulative sum and squared probability errors are also calculated for climatology using an even, constant probability bin width that is equal to 1 divided by the number of ensemble members (ex. three ensemble members are evaluated with three climatology probability bins of width 0.333). The accompanying skill score is derived by comparing the forecast RPS to the reference climatology RPS. Ensemble Observation Error -------------------------- In an attempt to ameliorate the effect of observation errors on the verification of forecasts, a random perturbation approach has been implemented. A great deal of user flexibility has been built in, but the methods detailed in :ref:`Candille and Talagrand (2008) ` can be replicated using the appropriate options. Additional variations of the ignorance score that include observational uncertainty recommended by :ref:`Ferro, 2017 ` are also provided. -Observation error information can be defined directly in the Ensemble-Stat configuration file or through a more flexible observation error lookup table. The user selects a distribution for the observation error, along with parameters for that distribution. Rescaling and bias correction can also be specified prior to the perturbation. Random draws from the distribution can then be added to either, or both of the forecast and observed fields, including ensemble members. Details about the effects of the choices on verification statistics should be considered, with many details provided in the literature (*e.g.* :ref:`Candille and Talagrand, 2008 `; :ref:`Saetra et al., 2004 `; :ref:`Santos and Ghelli, 2012 `). Generally, perturbation makes verification statistics better when applied to ensemble members, and worse when applied to the observations themselves. +Observation error information can be defined directly in the Ensemble-Stat configuration file or through a more flexible observation error lookup table. The user selects a distribution for the observation error, along with parameters for that distribution. Rescaling and bias correction can also be specified prior to the perturbation. Random draws from the distribution can then be added to either, or both of the forecast and observed fields, including ensemble members. Details about the effects of the choices on verification statistics should be considered, with many details provided in the literature (*e.g.*, :ref:`Candille and Talagrand, 2008 `; :ref:`Saetra et al., 2004 `; :ref:`Santos and Ghelli, 2012 `). Generally, perturbation makes verification statistics better when applied to ensemble members, and worse when applied to the observations themselves. -Normal and uniform are common choices for the observation error distribution. The uniform distribution provides the benefit of being bounded on both sides, thus preventing the perturbation from taking on extreme values. Normal is the most common choice for observation error. However, the user should realize that with the very large samples typical in NWP, some large outliers will almost certainly be introduced with the perturbation. For variables that are bounded below by 0, and that may have inconsistent observation errors (e.g. larger errors with larger measurements), a lognormal distribution may be selected. Wind speeds and precipitation measurements are the most common of this type of NWP variable. The lognormal error perturbation prevents measurements of 0 from being perturbed, and applies larger perturbations when measurements are larger. This is often the desired behavior in these cases, but this distribution can also lead to some outliers being introduced in the perturbation step. +Normal and uniform are common choices for the observation error distribution. The uniform distribution provides the benefit of being bounded on both sides, thus preventing the perturbation from taking on extreme values. Normal is the most common choice for observation error. However, the user should realize that with the very large samples typical in NWP, some large outliers will almost certainly be introduced with the perturbation. For variables that are bounded below by 0, and that may have inconsistent observation errors (e.g., larger errors with larger measurements), a lognormal distribution may be selected. Wind speeds and precipitation measurements are the most common of this type of NWP variable. The lognormal error perturbation prevents measurements of 0 from being perturbed, and applies larger perturbations when measurements are larger. This is often the desired behavior in these cases, but this distribution can also lead to some outliers being introduced in the perturbation step. Observation errors differ according to instrument, temporal and spatial representation, and variable type. Unfortunately, many observation errors have not been examined or documented in the literature. Those that have usually lack information regarding their distributions and approximate parameters. Instead, a range or typical value of observation error is often reported and these are often used as an estimate of the standard deviation of some distribution. Where possible, it is recommended to use the appropriate type and size of perturbation for the observation to prevent spurious results. @@ -84,7 +84,7 @@ This section contains information about configuring and running the Ensemble-Sta ensemble_stat Usage ------------------- -The usage statement for the Ensemble Stat tool is shown below: +The usage statement for the Ensemble-Stat tool is shown below: .. code-block:: none @@ -104,8 +104,8 @@ The usage statement for the Ensemble Stat tool is shown below: ensemble_stat has two required arguments and accepts several optional ones. -Required Arguments ensemble_stat -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +Required Arguments for ensemble_stat +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ 1. The **n_ens file_1 ... file_n | file_list** specifies either the number of ensemble members followed by a list of ensemble member file names or an ASCII file list of the file names to be used, as described in :numref:`ascii_file_lists`. @@ -138,21 +138,21 @@ An example of the ensemble_stat calling sequence is shown below: .. code-block:: none - ensemble_stat \ - 6 sample_fcst/2009123112/*gep*/d01_2009123112_02400.grib \ - config/EnsembleStatConfig \ - -grid_obs sample_obs/ST4/ST4.2010010112.24h \ - -point_obs out/ascii2nc/precip24_2010010112.nc \ - -outdir out/ensemble_stat -v 2 + ensemble_stat \ + 6 sample_fcst/2009123112/*gep*/d01_2009123112_02400.grib \ + config/EnsembleStatConfig \ + -grid_obs sample_obs/ST4/ST4.2010010112.24h \ + -point_obs out/ascii2nc/precip24_2010010112.nc \ + -outdir out/ensemble_stat -v 2 -In this example, the Ensemble-Stat tool will process six forecast files specified in the file list into an ensemble forecast. Observations in both point and grid format will be included, and be used to compute ensemble statistics separately. Ensemble Stat will create a NetCDF file containing requested ensemble fields and an output STAT file. +In this example, the Ensemble-Stat tool will process six forecast files specified in the file list into an ensemble forecast. Observations in both point and grid format will be included, and be used to compute ensemble statistics separately. Ensemble-Stat will create a NetCDF file containing requested ensemble fields and an output STAT file. ensemble_stat Configuration File -------------------------------- The default configuration file for the Ensemble-Stat tool named **EnsembleStatConfig_default** can be found in the installed *share/met/config* directory. Another version is located in *scripts/config*. We encourage users to make a copy of these files prior to modifying their contents. Each configuration file (both the default and sample) contains many comments describing its contents. The contents of the configuration file are also described in the subsections below. -Note that environment variables may be used when editing configuration files, as described in the :numref:`config_env_vars`. +Note that environment variables may be used when editing configuration files, as described in :numref:`config_env_vars`. ____________________ @@ -211,7 +211,7 @@ When processing the **fcst** data, compute a ratio of the number of valid ensemb When processing the **fcst** data, for each grid point compute a ratio of the number of valid data values to the number of ensemble members. If that ratio is less than **vld_thresh**, write out bad data. This threshold must be between 0 and 1. Setting this threshold to 1 will require each grid point to contain valid data for all ensemble members. -For each **field** listed in the forecast field, give the name and vertical or accumulation level, plus one or more categorical thresholds. The thresholds are specified using symbols, as shown above. It is the user's responsibility to know the units for each model variable and to choose appropriate threshold values. The thresholds are used to define ensemble relative frequencies, e.g. a threshold of >=5 can be used to compute the proportion of ensemble members predicting precipitation of at least 5mm at each grid point. +For each **field** listed in the forecast field, give the name and vertical or accumulation level, plus one or more categorical thresholds. The thresholds are specified using symbols, as shown above. It is the user's responsibility to know the units for each model variable and to choose appropriate threshold values. The thresholds are used to define ensemble relative frequencies, e.g., a threshold of >=5 can be used to compute the proportion of ensemble members predicting precipitation of at least 5mm at each grid point. _______________________ @@ -302,7 +302,7 @@ ____________________ skip_const = FALSE; -Setting **skip_const** to true tells Ensemble-Stat to exclude pairs where all the ensemble members and the observation have a constant value. For example, exclude points with zero precipitation amounts from all output line types. This option may be set separately for each **obs.field** entry. When set to false, constant points are and the observation rank is chosen at random. +Setting **skip_const** to true tells Ensemble-Stat to exclude pairs where all the ensemble members and the observation have a constant value. For example, exclude points with zero precipitation amounts from all output line types. This option may be set separately for each **obs.field** entry. When set to false, constant points are included and the observation rank is chosen at random. ____________________ @@ -326,7 +326,7 @@ ____________________ prob_pct_thresh = []; -The **prob_cat_thresh** entry is an array of thresholds. It is applied both to the computation of the RPS line type as well as the when generating probabilistic output line types. Since these thresholds can change for each variable, they can be specified separately for each **fcst.field** entry. If left empty but climatological mean and standard deviation data is provided, the **climo_cdf** thresholds will be used instead. If no climatology data is provided, and the RPS output line type is requested, then the **prob_cat_thresh** array must be defined. When probabilistic output line types are requested, for each **prob_cat_thresh** threshold listed, ensemble relative frequencies are derived and verified against the point and/or gridded observations. +The **prob_cat_thresh** entry is an array of thresholds. It is applied both to the computation of the RPS line type as well as when generating probabilistic output line types. Since these thresholds can change for each variable, they can be specified separately for each **fcst.field** entry. If left empty but climatological mean and standard deviation data is provided, the **climo_cdf** thresholds will be used instead. If no climatology data is provided, and the RPS output line type is requested, then the **prob_cat_thresh** array must be defined. When probabilistic output line types are requested, for each **prob_cat_thresh** threshold listed, ensemble relative frequencies are derived and verified against the point and/or gridded observations. The **prob_pct_thresh** entry is an array of thresholds which define the Nx2 probabilistic contingency table used to evaluate probability forecasts. It can be specified separately for each **fcst.field** entry. Options for defining probability bins are discussed in Point-Stat section :numref:`PS_Probability`. @@ -354,7 +354,7 @@ The **obs_error** dictionary controls how observation error information should b The **flag** entry toggles the observation error logic on (**TRUE**) and off (**FALSE**). When the **flag** is **TRUE**, random observation error perturbations are applied to the ensemble member values. No perturbation is applied to the observation values but the bias scale and offset values, if specified, are applied. -The **dist_type** entry may be set to **NONE, NORMAL, LOGNORMAL, EXPONENTIAL,CHISQUARED, GAMMA, UNIFORM**, or **BETA**. The default value of **NONE** indicates that the observation error table file should be used rather than the configuration file settings. +The **dist_type** entry may be set to **NONE, NORMAL, LOGNORMAL, EXPONENTIAL, CHISQUARED, GAMMA, UNIFORM**, or **BETA**. The default value of **NONE** indicates that the observation error table file should be used rather than the configuration file settings. The **dist_parm** entry is an array of length 1 or 2 specifying the parameters for the distribution selected in **dist_type**. The **GAMMA, UNIFORM**, and **BETA** distributions are defined by two parameters, specified as a comma-separated list (a,b), whereas all other distributions are defined by a single parameter. @@ -446,10 +446,10 @@ __________________ .. code-block:: none - nc_var_str = ""; + nc_var_str = ""; -The **nc_var_str** entry specifies a string for each ensemble field and verification task. This string is parsed from each **ens.field** and **obs.field** dictionary entry and is used to customize the variable names written to theNetCDF output file. The default is an empty string, meaning that no customization is applied to the output variable names. When the Ensemble-Stat config file contains two fields with the same name and level value, this entry is used to make the resulting variable names unique. +The **nc_var_str** entry specifies a string for each ensemble field and verification task. This string is parsed from each **ens.field** and **obs.field** dictionary entry and is used to customize the variable names written to the NetCDF output file. The default is an empty string, meaning that no customization is applied to the output variable names. When the Ensemble-Stat config file contains two fields with the same name and level value, this entry is used to make the resulting variable names unique. ________________ @@ -485,7 +485,7 @@ ensemble_stat_PREFIX_YYYYMMDD_HHMMSSV.stat where PREFIX indicates the user-defin The output ASCII files are named similarly: -ensemble_stat_PREFIX_YYYYMMDD_HHMMSSV_TYPE.txt where TYPE is one of elements of the **output_flag** configuration option to indicate the line type it contains. +ensemble_stat_PREFIX_YYYYMMDD_HHMMSSV_TYPE.txt where TYPE is one of the elements of the **output_flag** configuration option to indicate the line type it contains. When verification against gridded analyses is performed, Ensemble-Stat can produce output NetCDF files using the following naming convention: @@ -514,7 +514,7 @@ Spread/Skill Variance Ensemble Matched Pair information -The format of the STAT and ASCII output of the Ensemble-Stat tool are described below. +The format of the STAT and ASCII output of the Ensemble-Stat tool is described below. .. _table_ES_header_info_es_out: @@ -671,15 +671,15 @@ The format of the STAT and ASCII output of the Ensemble-Stat tool are described - Double * - 33 - ME_OERR - - The Mean Error of the PERTURBED ensemble mean (e.g. with Observation Error) + - The Mean Error of the PERTURBED ensemble mean (e.g., with Observation Error) - Double * - 34 - RMSE_OERR - - The Root Mean Square Error of the PERTURBED ensemble mean (e.g. with Observation Error) + - The Root Mean Square Error of the PERTURBED ensemble mean (e.g., with Observation Error) - Double * - 35 - SPREAD_OERR - - The square root of the mean of the variance of the PERTURBED ensemble member values (e.g. with Observation Error) at each observation location + - The square root of the mean of the variance of the PERTURBED ensemble member values (e.g., with Observation Error) at each observation location - Double * - 36 - SPREAD_PLUS_OERR @@ -715,7 +715,7 @@ The format of the STAT and ASCII output of the Ensemble-Stat tool are described - Double * - 44 - MAE_OERR - - The Mean Absolute Error of the PERTURBED ensemble mean (e.g. with Observation Error) + - The Mean Absolute Error of the PERTURBED ensemble mean (e.g., with Observation Error) - Double * - 45 - BIAS_RATIO @@ -735,15 +735,15 @@ The format of the STAT and ASCII output of the Ensemble-Stat tool are described - Integer * - 49 - ME_LT_OBS - - The Mean Error of the ensemble values less than or equal to their observations + - The Mean Error of the ensemble values less than their observations - Double * - 50 - IGN_CONV_OERR - - Error-convolved logarithmic scoring rule (i.e. ignornance score) from Equation 5 of :ref:`Ferro, 2017 ` + - Error-convolved logarithmic scoring rule (i.e., ignorance score) from Equation 5 of :ref:`Ferro, 2017 ` - Double * - 51 - IGN_CORR_OERR - - Error-corrected logarithmic scoring rule (i.e. ignornance score) from Equation 7 of :ref:`Ferro, 2017 ` + - Error-corrected logarithmic scoring rule (i.e., ignorance score) from Equation 7 of :ref:`Ferro, 2017 ` - Double .. _table_ES_header_info_es_out_RPS: @@ -766,7 +766,7 @@ The format of the STAT and ASCII output of the Ensemble-Stat tool are described - Integer * - 26 - N_PROB - - Number of probability thresholds (i.e. number of ensemble members in Ensemble-Stat) + - Number of probability thresholds (i.e., number of ensemble members in Ensemble-Stat) - Integer * - 27 - RPS_REL @@ -961,12 +961,12 @@ The format of the STAT and ASCII output of the Ensemble-Stat tool are described - The spread (standard deviation) of the unperturbed ensemble member values - Double * - Last-5 - - ENS_MEAN _OERR - - The PERTURBED ensemble mean (e.g. with Observation Error) + - ENS_MEAN_OERR + - The PERTURBED ensemble mean (e.g., with Observation Error) - Double * - Last-4 - SPREAD_OERR - - The spread (standard deviation) of the PERTURBED ensemble member values (e.g. with Observation Error) + - The spread (standard deviation) of the PERTURBED ensemble member values (e.g., with Observation Error) - Double * - Last-3 - SPREAD_PLUS_OERR @@ -986,7 +986,7 @@ The format of the STAT and ASCII output of the Ensemble-Stat tool are described - Double .. role:: raw-html(raw) - :format: html + :format: html .. _table_ES_header_info_es_out_SSVAR: @@ -1056,7 +1056,7 @@ The format of the STAT and ASCII output of the Ensemble-Stat tool are described - Double * - 39-41 - FSTDEV, :raw-html:`
` FSTDEV_NCL, :raw-html:`
` FSTDEV_NCU - - Standard deviation of the error including normal upper and lower confidence limits + - Standard deviation of the forecasts including normal upper and lower confidence limits - Double * - 42-43 - OBAR_NCL, :raw-html:`
` OBAR_NCU @@ -1064,7 +1064,7 @@ The format of the STAT and ASCII output of the Ensemble-Stat tool are described - Double * - 44-46 - OSTDEV, :raw-html:`
` OSTDEV_NCL, :raw-html:`
` OSTDEV_NCU - - Standard deviation of the error including normal upper and lower confidence limits + - Standard deviation of the observations including normal upper and lower confidence limits - Double * - 47-49 - PR_CORR, :raw-html:`
` PR_CORR_NCL, :raw-html:`
` PR_CORR_NCU diff --git a/docs/Users_Guide/gen-ens-prod.rst b/docs/Users_Guide/gen-ens-prod.rst index 304791ab5a..c24f1a398c 100644 --- a/docs/Users_Guide/gen-ens-prod.rst +++ b/docs/Users_Guide/gen-ens-prod.rst @@ -32,9 +32,9 @@ The Gen-Ens-Prod tool writes the gridded relative frequencies, NEP, NMEP, and EA Climatology Data ---------------- -The ensemble relative frequencies derived by Gen-Ens-Prod are computed by applying threshold(s) to the input ensemble member data. Those thresholds can be simple and remain constant over the entire domain (e.g. >0) or can be defined relative to the climatological distribution at each grid point (e.g. >OCDP90, for exceeding the 90-th percentile of the observation climatology data provided). +The ensemble relative frequencies derived by Gen-Ens-Prod are computed by applying threshold(s) to the input ensemble member data. Those thresholds can be simple and remain constant over the entire domain (e.g., >0) or can be defined relative to the climatological distribution at each grid point (e.g., >OCDP90, for exceeding the 90-th percentile of the observation climatology data provided). -To use climatological distribution percentile thresholds, users must specify the climatological mean ("climo_mean") and standard deviation ("climo_stdev") entries in the configuration file. With forecast climatology inputs, use forecast climatology distribution percentile thresholds (e.g. >FCDP90). With observation climatology inputs, use observation climatological distribution percentile thresholds instead (e.g. >OCDP90). However, Gen-Ens-Prod cannot actually determine the input climatology data source and both "FCDP" and "OCDP" threshold types will work. +To use climatological distribution percentile thresholds, users must specify the climatological mean ("climo_mean") and standard deviation ("climo_stdev") entries in the configuration file. With forecast climatology inputs, use forecast climatology distribution percentile thresholds (e.g., >FCDP90). With observation climatology inputs, use observation climatological distribution percentile thresholds instead (e.g., >OCDP90). However, Gen-Ens-Prod cannot actually determine the input climatology data source and both "FCDP" and "OCDP" threshold types will work. Practical Information ===================== @@ -44,7 +44,7 @@ This section contains information about configuring and running the Gen-Ens-Prod gen_ens_prod Usage ------------------ -The usage statement for the Ensemble Stat tool is shown below: +The usage statement for the Gen-Ens-Prod tool is shown below: .. code-block:: none @@ -58,8 +58,8 @@ The usage statement for the Ensemble Stat tool is shown below: gen_ens_prod has three required arguments and accepts several optional ones. -Required Arguments gen_ens_prod -------------------------------- +Required Arguments for gen_ens_prod +----------------------------------- 1. The **-ens file_1 ... file_n | file_list** option specifies the ensemble member files or ASCII file list of file names to be used, as described in :numref:`ascii_file_lists`. @@ -80,10 +80,10 @@ An example of the gen_ens_prod calling sequence is shown below: .. code-block:: none - gen_ens_prod \ - -ens sample_fcst/2009123112/*gep*/d01_2009123112_02400.grib \ - -out out/gen_ens_prod/gen_ens_prod_20100101_120000V_ens.nc \ - -config config/GenEnsProdConfig -v 2 + gen_ens_prod \ + -ens sample_fcst/2009123112/*gep*/d01_2009123112_02400.grib \ + -out out/gen_ens_prod/gen_ens_prod_20100101_120000V_ens.nc \ + -config config/GenEnsProdConfig -v 2 In this example, the Gen-Ens-Prod tool derives products from the input ensemble members listed on the command line. @@ -92,7 +92,7 @@ gen_ens_prod Configuration File The default configuration file for the Gen-Ens-Prod tool named **GenEnsProdConfig_default** can be found in the installed *share/met/config* directory. Another version is located in *scripts/config*. We encourage users to make a copy of these files prior to modifying their contents. The contents of the configuration file are described in the subsections below. -Note that environment variables may be used when editing configuration files, as described in the :numref:`config_env_vars`. +Note that environment variables may be used when editing configuration files, as described in :numref:`config_env_vars`. ____________________ @@ -129,11 +129,11 @@ _____________________ The **ens** dictionary defines which ensemble fields should be processed. -When summarizing the ensemble, compute a ratio of the number of valid ensemble fields to the total number of ensemble members. If this ratio is less than the **ens_thresh**, then quit with an error. This threshold must be between 0 and 1. Setting this threshold to 1 requires that all ensemble members input files exist and all requested data be present. +When summarizing the ensemble, compute a ratio of the number of valid ensemble fields to the total number of ensemble members. If this ratio is less than the **ens_thresh**, then quit with an error. This threshold must be between 0 and 1. Setting this threshold to 1 requires that all ensemble member input files exist and all requested data be present. When summarizing the ensemble, for each grid point compute a ratio of the number of valid data values to the number of ensemble members. If that ratio is less than **vld_thresh**, write out bad data for that grid point. This threshold must be between 0 and 1. Setting this threshold to 1 requires that each grid point contain valid data for all ensemble members in order to compute ensemble product values for that grid point. -For each dictionary entry in the **field** array, give the name and vertical or accumulation level, plus one or more categorical thresholds in the **cat_thresh** entry. The formatting for threshold are described in :numref:`config_options`. It is the user's responsibility to know the units for each model variable and choose appropriate threshold values. The thresholds are used to define ensemble relative frequencies. For example, a threshold of >=5 is used to define the proportion of ensemble members predicting precipitation of at least 5mm at each grid point. +For each dictionary entry in the **field** array, give the name and vertical or accumulation level, plus one or more categorical thresholds in the **cat_thresh** entry. The formatting for thresholds is described in :numref:`config_options`. It is the user's responsibility to know the units for each model variable and choose appropriate threshold values. The thresholds are used to define ensemble relative frequencies. For example, a threshold of >=5 is used to define the proportion of ensemble members predicting precipitation of at least 5mm at each grid point. _______________________ @@ -197,7 +197,7 @@ _____________________ normalize = NONE; -The **normalize** option defines if and how the input ensemble member data should be normalized. Options are provided to normalize relative to an external climatology, specified using the **climo_mean** and **climo_stdev** dictionaries, or relative to current ensemble forecast being processed. The anomaly is computed by subtracting the (climatological or ensemble) mean from each ensemble memeber. The standard anomaly is computed by dividing the anomaly by the (climatological or ensemble) standard deviation. Values for the **normalize** option are described below: +The **normalize** option defines if and how the input ensemble member data should be normalized. Options are provided to normalize relative to an external climatology, specified using the **climo_mean** and **climo_stdev** dictionaries, or relative to current ensemble forecast being processed. The anomaly is computed by subtracting the (climatological or ensemble) mean from each ensemble member. The standard anomaly is computed by dividing the anomaly by the (climatological or ensemble) standard deviation. Values for the **normalize** option are described below: • **NONE** (default) to skip the normalization step and process the raw ensemble member data. @@ -246,7 +246,7 @@ _____________________ Similar to the **interp** dictionary, the **nmep_smooth** dictionary includes a **type** array of dictionaries to define one or more methods for smoothing the NMEP data. Setting the interpolation method to nearest neighbor (**NEAREST**) effectively disables this smoothing step. -If **ensemble_flag.nmep** is set to TRUE, NMEP output is created for each combination of the categorical threshold (**cat_thresh**), neighborhood width (**nbrhd_prob.width**), and smoothing method(**nmep_smooth.type**) specified. +If **ensemble_flag.nmep** is set to TRUE, NMEP output is created for each combination of the categorical threshold (**cat_thresh**), neighborhood width (**nbrhd_prob.width**), and smoothing method (**nmep_smooth.type**) specified. _____________________ @@ -263,7 +263,7 @@ _____________________ The **eas_prob** dictionary defines the options for the Ensemble Agreement Scale (EAS) probability method. -The **shape** is a **SQUARE** or **CIRCLE** centered on the current point, and the **width** array specifies the candidate widths of the square or diameter of the circle as an odd integer. The **vld_thresh** entry is a number between 0 and 1 specifying the required ratio of valid data in the neighborhood for an output value to be computed. **alpha** is a number between 0 and 1 that defines the EAS distance criteria. **guassian_dx** and **gaussian_radius** define the Gaussian smoother that is applied to the raw EAS probability values. +The **shape** is a **SQUARE** or **CIRCLE** centered on the current point, and the **width** array specifies the candidate widths of the square or diameter of the circle as an odd integer. The **vld_thresh** entry is a number between 0 and 1 specifying the required ratio of valid data in the neighborhood for an output value to be computed. **alpha** is a number between 0 and 1 that defines the EAS distance criteria. **gaussian_dx** and **gaussian_radius** define the Gaussian smoother that is applied to the raw EAS probability values. If **ensemble_flag.eas** or **ensemble_flag.eas_width** is set to TRUE, the EAS algorithm is run for each categorical threshold (**cat_thresh**) specified. The **eas** and **eas_width** flags control the writing of the EAS probabilities and widths chosen, respectively. @@ -310,7 +310,7 @@ The **ensemble_flag** specifies which derived ensemble fields should be calculat 9. Ensemble Valid Data Count -10. Ensemble Relative Frequency (i.e. uncalibrate probability forecast) for each categorical threshold (**cat_thresh**) specified +10. Ensemble Relative Frequency (i.e., uncalibrated probability forecast) for each categorical threshold (**cat_thresh**) specified 11. Neighborhood Ensemble Probability for each categorical threshold (**cat_thresh**) and neighborhood width (**nbrhd_prob.width**) specified @@ -325,6 +325,6 @@ The **ensemble_flag** specifies which derived ensemble fields should be calculat gen_ens_prod Output ------------------- -The Gen-Ens-Prod tools writes a gridded NetCDF output file whose file name is specified using the -out command line option. The contents of that file depend on the contents of the **ens.field** array, the **ensemble_flag** options selected, and the presence of climatology data. The NetCDF variable names are self-describing and include the name/level of the field being processed, the type of ensemble product, and any relevant threshold information. If **nc_var_str** is defined for an **ens.field** array entry, that string is included in the corresponding NetCDF output variable names. +The Gen-Ens-Prod tool writes a gridded NetCDF output file whose file name is specified using the -out command line option. The contents of that file depend on the contents of the **ens.field** array, the **ensemble_flag** options selected, and the presence of climatology data. The NetCDF variable names are self-describing and include the name/level of the field being processed, the type of ensemble product, and any relevant threshold information. If **nc_var_str** is defined for an **ens.field** array entry, that string is included in the corresponding NetCDF output variable names. -The Gen-Ens-Prod NetCDF output can be passed as input to the MET statistics tools, like Point-Stat and Grid-Stat, for futher processing and comparison against observations. +The Gen-Ens-Prod NetCDF output can be passed as input to the MET statistics tools, like Point-Stat and Grid-Stat, for further processing and comparison against observations. diff --git a/docs/Users_Guide/grid-diag.rst b/docs/Users_Guide/grid-diag.rst index a53834a4ad..63084c2691 100644 --- a/docs/Users_Guide/grid-diag.rst +++ b/docs/Users_Guide/grid-diag.rst @@ -91,7 +91,7 @@ _____________________ The **power_spectrum** dictionary defines options for computing power spectra and can be specified separately for each **data.field** entry below. -The **missing_flag** and **missing_value** entries define how bad data values should be handled. For all other output types, bad data values are ignored but they are problematic for power spectra. Set **missing_flag** to **NONE** (default) to skip power spectrum when bad data is present, to **MEAN** to replace bad data with the mean of each input field, or to **VALUE** to replace bad data with the constant numeric value specified by **missing_value**. Set **vld_thresh** to a number beween 0 and 1 to define the required ratio of valid data to be present to compute power spectra output for that field. +The **missing_flag** and **missing_value** entries define how bad data values should be handled. For all other output types, bad data values are ignored but they are problematic for power spectra. Set **missing_flag** to **NONE** (default) to skip power spectrum when bad data is present, to **MEAN** to replace bad data with the mean of each input field, or to **VALUE** to replace bad data with the constant numeric value specified by **missing_value**. Set **vld_thresh** to a number between 0 and 1 to define the required ratio of valid data to be present to compute power spectra output for that field. _____________________ @@ -122,18 +122,18 @@ _____________________ .. code-block:: none - output_flag = { - histogram_1d = TRUE; - histogram_2d = TRUE; - info_theory = FALSE; - power_spectrum = FALSE; - } + output_flag = { + histogram_1d = TRUE; + histogram_2d = TRUE; + info_theory = FALSE; + power_spectrum = FALSE; + } The **output_flag** dictionary controls the type of output that the Grid-Diag tool generates. Each flag should be set to **TRUE** or **FALSE** to enable the computation and writing of one or more variables to the output NetCDF file, as described below: -1. **histogram_1d** for 1-dimensional histograms for each **data.field** entry, including minimum, maxmimum, and midpoint values for each histogram bin. +1. **histogram_1d** for 1-dimensional histograms for each **data.field** entry, including minimum, maximum, and midpoint values for each histogram bin. -2. **histogram_2d** for 2-dimensional histograms for each pair of **data.field** entries, including minimum, maxmimum, and midpoint values for each histogram bin. +2. **histogram_2d** for 2-dimensional histograms for each pair of **data.field** entries, including minimum, maximum, and midpoint values for each histogram bin. 3. **info_theory** for information theory metrics, including entropy for each **data.field** entry and mutual information and joint entropy for each pair of entries. @@ -144,13 +144,13 @@ grid_diag Output File The output NetCDF file contains variables for **grid_size** and **n_series** which specify the number of points in the grid and the number of files that were processed, respectively. The range of the initialization, valid, and lead times processed is written to the global attributes. These variables and global attributes are written for each run. -If histogram or information theory output is requested, dimensions are created for the number of masking regions and one for each of the specified data variable and level combinations, e.g. APCP_L0 and PWAT_L0. The bin minimum and maximum values are indicated with an _min or _max appended to the variable/level. For each variable and level combination, a coordinate variable is written to indicate the midpoint value for each histogram bin. +If histogram or information theory output is requested, dimensions are created for the number of masking regions and one for each of the specified data variable and level combinations, e.g., APCP_L0 and PWAT_L0. The bin minimum and maximum values are indicated with an _min or _max appended to the variable/level. For each variable and level combination, a coordinate variable is written to indicate the midpoint value for each histogram bin. The **mask_name** and **mask_size** variables have dimensions based on the number of masking regions and indicate the name of each masking region and the number of grid points it includes, respectively. Masking variables are written when histogram or information theory output is requested whereas power spectrum output is computed over the full input grid. If 1-dimensional histograms are requested, a corresponding **hist_** variable is written for each variable/level in the data dictionary. This variable has dimensions for the number of masking regions and for the number of bins specified in the data dictionary. For example, hist_APCP_L0 and hist_PWAT_L0 are the counts of all data values falling within each bin for a given spatial masking region. Data values below the minimum or above the maximum are included in the lowest and highest bins, respectively. A warning message is printed when the range of the data falls outside the range defined in the configuration file. In this case, users are advised to adjust the **range** setting and rerun. -If 2-dimensional joint historgrams are requested, a corresponding **hist_** variable is written for each combination of variable/level entries in the data dictionary. This variable has dimensions for the number of masking regions and for the number of bins specified for the two data dictionary entries. For example, hist_APCP_L0_PWAT_L0 is the joint histogram for those two variables/levels for a given spatial masking region. +If 2-dimensional joint histograms are requested, a corresponding **hist_** variable is written for each combination of variable/level entries in the data dictionary. This variable has dimensions for the number of masking regions and for the number of bins specified for the two data dictionary entries. For example, hist_APCP_L0_PWAT_L0 is the joint histogram for those two variables/levels for a given spatial masking region. If information theory output is requested, **entropy_**, **joint_entropy_**, and **mutual_information_** variables are written. Shannon entropy is derived from each 1-dimensional histogram, while joint entropy and mutual information are derived from each 2-dimensional joint histogram. These variables have one dimension for the number of masking regions and are computed using log base 2 rather than the natural logarithm. As such, their units are specified in the output as "bits" rather than "nats". diff --git a/docs/Users_Guide/grid-stat.rst b/docs/Users_Guide/grid-stat.rst index ec0c013a27..e5a7527649 100644 --- a/docs/Users_Guide/grid-stat.rst +++ b/docs/Users_Guide/grid-stat.rst @@ -31,7 +31,7 @@ Measures for Continuous Variables For continuous variables, many verification measures are based on the forecast error (i.e., f - o). However, it also is of interest to investigate characteristics of the forecasts, and the observations, as well as their relationship. These concepts are consistent with the general framework for verification outlined by :ref:`Murphy and Winkler (1987) `. The statistics produced by MET for continuous forecasts represent this philosophy of verification, which focuses on a variety of aspects of performance rather than a single measure. See :numref:`Appendix C, Section %s ` for specific information. -A user may wish to eliminate certain values of the forecasts from the calculation of statistics, a process referred to here as "conditional verification". For example, a user may eliminate all temperatures above freezing and then calculate the error statistics only for those forecasts of below freezing temperatures. Another common example involves verification of wind forecasts. Since wind direction is indeterminate at very low wind speeds, the user may wish to set a minimum wind speed threshold prior to calculating error statistics for wind direction. The user may specify these thresholds in the configuration file to specify the conditional verification. Thresholds can be specified using the usual Fortran conventions (<, <=, ==, !-, >=, or >) followed by a numeric value. The threshold type may also be specified using two letter abbreviations (lt, le, eq, ne, ge, gt). Further, more complex thresholds can be achieved by defining multiple thresholds and using && or || to string together event definition logic. The forecast and observation threshold can be used together according to user preference by specifying one of: UNION, INTERSECTION, or SYMDIFF (symmetric difference). +A user may wish to eliminate certain values of the forecasts from the calculation of statistics, a process referred to here as "conditional verification". For example, a user may eliminate all temperatures above freezing and then calculate the error statistics only for those forecasts of below freezing temperatures. Another common example involves verification of wind forecasts. Since wind direction is indeterminate at very low wind speeds, the user may wish to set a minimum wind speed threshold prior to calculating error statistics for wind direction. The user may specify these thresholds in the configuration file to specify the conditional verification. Thresholds can be specified using the usual Fortran conventions (<, <=, ==, !=, >=, or >) followed by a numeric value. The threshold type may also be specified using two letter abbreviations (lt, le, eq, ne, ge, gt). Further, more complex thresholds can be achieved by defining multiple thresholds and using && or || to string together event definition logic. The forecast and observation threshold can be used together according to user preference by specifying one of: UNION, INTERSECTION, or SYMDIFF (symmetric difference). Measures for Probabilistic Forecasts and Dichotomous Outcomes ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ @@ -93,10 +93,10 @@ Gradient Statistics The S1 score has been in historical use for verification of forecasts, particularly for variables such as pressure and geopotential height. This score compares differences between adjacent grid points in the forecast and observed fields. When the adjacent points in both forecast and observed fields exhibit the same differences, the S1 score will be the perfect value of 0. Larger differences will result in a larger score. -Differences are computed in both of the horizontal grid directions and is not a true mathematical gradient. Because the S1 score focuses on differences only, any bias in the forecast will not be measured. Further, the score depends on the domain and spacing of the grid, so can only be compared on forecasts with identical grids. +Differences are computed in both of the horizontal grid directions and are not a true mathematical gradient. Because the S1 score focuses on differences only, any bias in the forecast will not be measured. Further, the score depends on the domain and spacing of the grid, so can only be compared on forecasts with identical grids. As described in :ref:`Ebert-Uphoff et al., 2024 `, statistics based -on the magnitude of the forecast and observed gradients are also provided. Similiar to +on the magnitude of the forecast and observed gradients are also provided. Similar to the S1 score, the root-mean-squared error of the magnitude of the gradients and their divergence quantify the similarity in the texture of the fields, with 0 being a perfect score. These gradient-based statistics assess the difference in smoothness between the @@ -113,13 +113,13 @@ Because these methods rely on the distance map, it is helpful to understand prec .. figure:: figure/grid-stat_fig1.png - The above diagram depicts how a distance map is formed. From every grid point in the domain (depicted by the larger rectangle), the shortest distance from that grid to the nearest non-zero grid point (event; depicted by the gray rectangle labeled as A) is calculated (a sample of grid points with arrows indicate the path of the shortest distance with the length of the arrow equal to this distance. In a distance map, the value at each grid point is this distance. For example, grid points within the rectangle A will all have value zero in the distance map. + The above diagram depicts how a distance map is formed. From every grid point in the domain (depicted by the larger rectangle), the shortest distance from that grid to the nearest non-zero grid point (event; depicted by the gray rectangle labeled as A) is calculated (a sample of grid points with arrows indicate the path of the shortest distance with the length of the arrow equal to this distance). In a distance map, the value at each grid point is this distance. For example, grid points within the rectangle A will all have value zero in the distance map. .. _grid-stat_fig2: .. figure:: figure/grid-stat_fig2.png - Diagram depicting the shortest distances from one event area to another. The yellow bar indicates the part of the event area A to where all of the shortest distances from B are calculated. That is, the shortest distances from every point inside the set B to the set A all point to a point along the yellow bar. + Diagram depicting the shortest distances from one event area to another. The yellow bar indicates the part of the event area A to where all of the shortest distances from B are calculated. That is, the shortest distances from every point inside the set B to the set A all point to a point along the yellow bar. While :numref:`grid-stat_fig1` and :numref:`grid-stat_fig2` are helpful in illustrating the idea of a distance map, :numref:`grid-stat_fig3` shows an actual distance map calculated for binary fields consisting of circular event areas, where one field has two circular event areas labeled A, and the second has one circular event area labeled B. Notice that the values of the distance map inside the event areas are all zero (dark blue) and the distances grow larger in the pattern of concentric circles around these event areas as grid cells move further away. Finally, :numref:`grid-stat_fig4` depicts special situations from which the distance map measures to be discussed are calculated. In particular, the top left panel shows the absolute difference between the two distance maps presented in the bottom row of :numref:`grid-stat_fig3`. The top right panel shows the portion of the distance map for A that falls within the event area of B, and the bottom left depicts the portion of the distance map for B that falls within the event area A. That is, the first shows the shortest distances from every grid point in the set B to the nearest grid point in the event area A, and the latter shows the shortest distance from every grid point in A to the nearest grid point in B. @@ -127,13 +127,13 @@ While :numref:`grid-stat_fig1` and :numref:`grid-stat_fig2` are helpful in illus .. figure:: figure/grid-stat_fig3.png - Binary fields (top) with event areas A (consisting of two circular event areas) and a second field with event area B (single circular area) with their respective distance maps (bottom). + Binary fields (top) with event areas A (consisting of two circular event areas) and a second field with event area B (single circular area) with their respective distance maps (bottom). .. _grid-stat_fig4: .. figure:: figure/grid-stat_fig4.png - The absolute difference between the distance maps in the bottom row of :numref:`grid-stat_fig3` (top left), the shortest distances from every grid point in B to the nearest grid point in A (top right), and the shortest distances from every grid point in A to the nearest grid points in B (bottom left). The latter two do not have axes in order to emphasize that the distances are now only considered from within the respective event sets. The top right graphic is the distance map of A conditioned on the presence of an event from B, and that in the bottom left is the distance map of B conditioned on the presence of an event from A. + The absolute difference between the distance maps in the bottom row of :numref:`grid-stat_fig3` (top left), the shortest distances from every grid point in B to the nearest grid point in A (top right), and the shortest distances from every grid point in A to the nearest grid points in B (bottom left). The latter two do not have axes in order to emphasize that the distances are now only considered from within the respective event sets. The top right graphic is the distance map of A conditioned on the presence of an event from B, and that in the bottom left is the distance map of B conditioned on the presence of an event from A. The statistics derived from these distance maps are described in :numref:`Appendix C, Section %s `. To make fair comparisons, any grid point containing bad data in either the forecast or observation field is set to bad data in both fields. For each combination of input field and categorical threshold requested in the configuration file, Grid-Stat applies that threshold to define events in the forecast and observation fields and computes distance maps for those binary fields. Statistics for all requested masking regions are derived from those distance maps. Note that the distance maps are computed only once over the full verification domain, not separately for each masking region. Events occurring outside of a masking region can affect the distance map values inside that masking region and, therefore, can also affect the distance maps statistics for that region. @@ -154,7 +154,7 @@ Whether or not the forecast from :numref:`grid-stat_fig6` is “good” or not d .. figure:: figure/grid-stat_fig6.png - Top left is an example of an accumulated precipitation (mm/h) forecast with the corresponding observed field on the top right. Bottom left shows the difference in binary fields, where the binary fields are created by setting all values in the original fields that fall above :math:`2.1 mmh^{-1}` to one and the rest to zero. Bottom right shows the results for :math:`G_\beta` calculated on the binary fields using the threshold of :math:`2.1 mmh^{-1}` over a range of choices for :math:`\beta`. + Top left is an example of an accumulated precipitation (mm/h) forecast with the corresponding observed field on the top right. Bottom left shows the difference in binary fields, where the binary fields are created by setting all values in the original fields that fall above :math:`2.1 mmh^{-1}` to one and the rest to zero. Bottom right shows the results for :math:`G_\beta` calculated on the binary fields using the threshold of :math:`2.1 mmh^{-1}` over a range of choices for :math:`\beta`. In some cases, a user may be interested in a much higher threshold than :math:`2.1 mmh^{-1}` of the above example. :ref:`Gilleland, 2021 (Fig. 4) `, for example, shows this same forecast using a threshold of :math:`40 mmh^{-1}`. Only a small area in Mississippi has such extreme rain predicted at this valid time; yet none was observed. Small spatial areas of extreme rain in the observed field, however, did occur in a location far away from Mississippi that was not predicted. Generally, for this type of verification, the Hausdorff metric is a good choice of measure. However, a small choice of :math:`\beta` will provide similar results as the Hausdorff distance (:ref:`Gilleland, 2021 `). The user should think about the average size of storm areas and multiply this value by the displacement distance they are comfortable with in order to get a good initial choice for :math:`\beta`, and may have to increase or decrease its value by trial-and-error using one or two example cases from their verification set. @@ -167,7 +167,7 @@ Practical Information This section contains information about configuring and running the Grid-Stat tool. The Grid-Stat tool verifies gridded model data using gridded observations. The input gridded model and observation datasets must be in one of the MET supported file formats. The requirement of having all gridded fields using the same grid specification was removed in METv5.1. There is a regrid option in the configuration file that allows the user to define the grid upon which the scores will be computed. The gridded observation data may be a gridded analysis based on observations such as Stage II or Stage IV data for verifying accumulated precipitation, or a model analysis field may be used. -The Grid-Stat tool provides the capability of verifying one or more model variables/levels using multiple thresholds for each model variable/level. The Grid-Stat tool performs no interpolation when the input model, observation, and climatology datasets must be on a common grid. MET will interpolate these files to a common grid if one is specified. The interpolation parameters may be used to perform a smoothing operation on the forecast field prior to verifying it to investigate how the scale of the forecast affects the verification statistics. The Grid-Stat tool computes a number of continuous statistics for the forecast minus observation differences, discrete statistics once the data have been thresholded, or statistics for probabilistic forecasts. All types of statistics can incorporate a climatological reference. +The Grid-Stat tool provides the capability of verifying one or more model variables/levels using multiple thresholds for each model variable/level. The Grid-Stat tool performs no interpolation when the input model, observation, and climatology datasets are already on a common grid. MET will interpolate these files to a common grid if one is specified. The interpolation parameters may be used to perform a smoothing operation on the forecast field prior to verifying it to investigate how the scale of the forecast affects the verification statistics. The Grid-Stat tool computes a number of continuous statistics for the forecast minus observation differences, discrete statistics once the data have been thresholded, or statistics for probabilistic forecasts. All types of statistics can incorporate a climatological reference. grid_stat Usage --------------- @@ -228,8 +228,8 @@ A second example of the grid_stat calling sequence is listed below: .. code-block:: none - grid_stat sample_fcst.nc - sample_obs.nc + grid_stat sample_fcst.nc \ + sample_obs.nc \ GridStatConfig In the second example, the Grid-Stat tool will verify the model data in the sample_fcst.nc NetCDF output of pcp_combine, using the observations in the sample_obs.nc NetCDF output of pcp_combine, and applying the configuration options specified in the **GridStatConfig** file. Because the model and observation files contain only a single field of accumulated precipitation, the **GridStatConfig** file should be configured to specify that only accumulated precipitation be verified. @@ -241,7 +241,7 @@ grid_stat Configuration File The default configuration file for the Grid-Stat tool, named **GridStatConfig_default**, can be found in the installed *share/met/config* directory. Other versions of the configuration file are included in *scripts/config*. We recommend that users make a copy of the default (or other) configuration file prior to modifying it. The contents are described in more detail below. -Note that environment variables may be used when editing configuration files, as described in the :numref:`config_env_vars`. +Note that environment variables may be used when editing configuration files, as described in :numref:`config_env_vars`. __________________________ @@ -316,7 +316,7 @@ ___________________ } -The **fourier** entry is a dictionary which specifies the application of the Fourier decomposition method. It consists of two arrays of the same length which define the beginning and ending wave numbers to be included. If the arrays have length zero, no Fourier decomposition is applied. For each array entry, the requested Fourier decomposition is applied to the forecast and observation fields. The beginning and ending wave numbers are indicated in the MET ASCII output files by the INTERP_MTHD column (e.g. WV1_0-3 for waves 0 to 3 or WV1_10 for only wave 10). This 1-dimensional Fourier decomposition is computed along the Y-dimension only (i.e. the columns of data). It is applied to the forecast and observation fields as well as the climatological mean field, if specified. It is only defined when each grid point contains valid data. If any input field contains missing data, no Fourier decomposition is computed. +The **fourier** entry is a dictionary which specifies the application of the Fourier decomposition method. It consists of two arrays of the same length which define the beginning and ending wave numbers to be included. If the arrays have length zero, no Fourier decomposition is applied. For each array entry, the requested Fourier decomposition is applied to the forecast and observation fields. The beginning and ending wave numbers are indicated in the MET ASCII output files by the INTERP_MTHD column (e.g., WV1_0-3 for waves 0 to 3 or WV1_10 for only wave 10). This 1-dimensional Fourier decomposition is computed along the Y-dimension only (i.e., the columns of data). It is applied to the forecast and observation fields as well as the climatological mean field, if specified. It is only defined when each grid point contains valid data. If any input field contains missing data, no Fourier decomposition is computed. The available wave numbers start at 0 (the mean across each row of data) and end at (Nx+1)/2 (the finest level of detail), where Nx is the X-dimension of the verification grid: @@ -340,7 +340,7 @@ _____________________ beta_value(n) = n * n / 2.0; } -The **distance_map** entry is a dictionary containing options related to the distance map statistics in the **DMAP** output line type. The **baddeley_p** entry is an integer specifying the exponent used in the Lp-norm when computing the Baddeley :math:`\Delta` metric. The **baddeley_max_dist** entry is a floating point number specifying the maximum allowable distance for each distance map. Any distances larger than this number will be reset to this constant. A value of **NA** indicates that no maximum distance value should be used. The **fom_alpha** entry is a floating point number specifying the scaling constant to be used when computing Pratt's Figure of Merit. The **zhu_weight** specifies a value between 0 and 1 to define the importance of the RMSE of the binary fields (i.e. amount of overlap) versus the mean-error distance (MED). The default value of 0.5 gives equal weighting. This configuration option may be set separately in each **obs.field** entry. The **beta_value** entry is defined as a function of n, where n is the total number of grid points in the full verification domain containing valid data in both the forecast and observation fields. The resulting beta_value is used to compute the :math:`G_\beta` statistic. The default function, :math:`N^2 / 2`, is recommended in :ref:`Gilleland, 2021 ` but can be modified as needed. +The **distance_map** entry is a dictionary containing options related to the distance map statistics in the **DMAP** output line type. The **baddeley_p** entry is an integer specifying the exponent used in the Lp-norm when computing the Baddeley :math:`\Delta` metric. The **baddeley_max_dist** entry is a floating point number specifying the maximum allowable distance for each distance map. Any distances larger than this number will be reset to this constant. A value of **NA** indicates that no maximum distance value should be used. The **fom_alpha** entry is a floating point number specifying the scaling constant to be used when computing Pratt's Figure of Merit. The **zhu_weight** specifies a value between 0 and 1 to define the importance of the RMSE of the binary fields (i.e., amount of overlap) versus the mean-error distance (MED). The default value of 0.5 gives equal weighting. This configuration option may be set separately in each **obs.field** entry. The **beta_value** entry is defined as a function of n, where n is the total number of grid points in the full verification domain containing valid data in both the forecast and observation fields. The resulting beta_value is used to compute the :math:`G_\beta` statistic. The default function, :math:`N^2 / 2`, is recommended in :ref:`Gilleland, 2021 ` but can be modified as needed. _____________________ @@ -481,7 +481,7 @@ The output ASCII files are named similarly: grid_stat_PREFIX_HHMMSSL_YYYYMMDD_HHMMSSV_TYPE.txt where TYPE is one of fho, ctc, cts, mctc, mcts, cnt, sl1l2, vl1l2, vcnt, pct, pstd, pjc, prc, eclv, nbrctc, nbrcts, nbrcnt, dmap, or grad to indicate the line type it contains. -The format of the STAT and ASCII output of the Grid-Stat tool are the same as the format of the STAT and ASCII output of the Point-Stat tool with the exception of the five additional line types. Please refer to the tables in :numref:`point_stat-output` for a description of the common output STAT and optional ASCII file line types. The formats of the five additional line types for grid_stat are explained in the following tables. +The format of the STAT and ASCII output of the Grid-Stat tool is the same as the format of the STAT and ASCII output of the Point-Stat tool with the exception of the five additional line types. Please refer to the tables in :numref:`point_stat-output` for a description of the common output STAT and optional ASCII file line types. The formats of the five additional line types for grid_stat are explained in the following tables. .. _table_GS_header_info_gs_outputs: @@ -626,7 +626,7 @@ The format of the STAT and ASCII output of the Grid-Stat tool are the same as th - Double .. role:: raw-html(raw) - :format: html + :format: html .. _table_GS_format_info_NBRCTS: @@ -703,23 +703,23 @@ The format of the STAT and ASCII output of the Grid-Stat tool are the same as th - Logarithm of the Odds Ratio including normal and bootstrap upper and lower confidence limits - Double * - 90-94 - - ORSS, :raw-html:`
` ORSS _NCL, :raw-html:`
` ORSS _NCU, :raw-html:`
` ORSS _BCL, :raw-html:`
` ORSS _BCU + - ORSS, :raw-html:`
` ORSS_NCL, :raw-html:`
` ORSS_NCU, :raw-html:`
` ORSS_BCL, :raw-html:`
` ORSS_BCU - Odds Ratio Skill Score including normal and bootstrap upper and lower confidence limits - Double * - 95-99 - - EDS, :raw-html:`
` EDS _NCL, :raw-html:`
` EDS _NCU, :raw-html:`
` EDS _BCL, :raw-html:`
` EDS _BCU + - EDS, :raw-html:`
` EDS_NCL, :raw-html:`
` EDS_NCU, :raw-html:`
` EDS_BCL, :raw-html:`
` EDS_BCU - Extreme Dependency Score including normal and bootstrap upper and lower confidence limits - Double * - 100-104 - - SEDS, :raw-html:`
` SEDS _NCL, :raw-html:`
` SEDS _NCU, :raw-html:`
` SEDS _BCL SEDS _BCU + - SEDS, :raw-html:`
` SEDS_NCL, :raw-html:`
` SEDS_NCU, :raw-html:`
` SEDS_BCL, :raw-html:`
` SEDS_BCU - Symmetric Extreme Dependency Score including normal and bootstrap upper and lower confidence limits - Double * - 105-109 - - EDI, :raw-html:`
` EDI _NCL, :raw-html:`
` EDI _NCU, :raw-html:`
` EDI _BCL, :raw-html:`
` EDI _BCU + - EDI, :raw-html:`
` EDI_NCL, :raw-html:`
` EDI_NCU, :raw-html:`
` EDI_BCL, :raw-html:`
` EDI_BCU - Extreme Dependency Index including normal and bootstrap upper and lower confidence limits - Double * - 110-114 - - SEDI, :raw-html:`
` SEDI _NCL, :raw-html:`
` SEDI _NCU, :raw-html:`
` SEDI _BCL,SEDI _BCU + - SEDI, :raw-html:`
` SEDI_NCL, :raw-html:`
` SEDI_NCU, :raw-html:`
` SEDI_BCL, :raw-html:`
` SEDI_BCU - Symmetric Extremal Dependency Index including normal and bootstrap upper and lower confidence limits - Double * - 115-117 @@ -729,11 +729,11 @@ The format of the STAT and ASCII output of the Grid-Stat tool are the same as th .. role:: raw-html(raw) - :format: html + :format: html .. _table_GS_format_info_NBRCNT: -.. list-table:: Format information for NBRCNT(Neighborhood Continuous Statistics) output line type +.. list-table:: Format information for NBRCNT (Neighborhood Continuous Statistics) output line type :widths: auto :header-rows: 1 @@ -766,11 +766,11 @@ The format of the STAT and ASCII output of the Grid-Stat tool are the same as th - Uniform Fractions Skill Score including bootstrap upper and lower confidence limits - Double * - 38-40 - - F_RATE, :raw-html:`
` F_RATE _BCL, :raw-html:`
` F_RATE _BCU + - F_RATE, :raw-html:`
` F_RATE_BCL, :raw-html:`
` F_RATE_BCU - Forecast event frequency including bootstrap upper and lower confidence limits - Double * - 41-43 - - O_RATE, :raw-html:`
` O _RATE _BCL, :raw-html:`
` O _RATE _BCU + - O_RATE, :raw-html:`
` O_RATE_BCL, :raw-html:`
` O_RATE_BCU - Observed event frequency including bootstrap upper and lower confidence limits - Double @@ -834,7 +834,7 @@ The format of the STAT and ASCII output of the Grid-Stat tool are the same as th - Double * - 36 - OGMAG - - Magnitude of the observed gradient when the X and Y-directions are intrepreted as a vector + - Magnitude of the observed gradient when the X and Y-directions are interpreted as a vector - Double * - 37 - MAG_RMSE @@ -956,7 +956,7 @@ The format of the STAT and ASCII output of the Grid-Stat tool are the same as th - Beta value used to compute :math:`G_\beta` - Double -If requested using the **nc_pairs_flag** dictionary in the configuration file, a NetCDF file containing the matched pair and forecast minus observation difference fields for each combination of variable type/level and masking region applied will be generated. The contents of this file are determined by the contents of the nc_pairs_flag dictionary. The output NetCDF file is named similarly to the other output files: **grid_stat_PREFIX_ HHMMSSL_YYYYMMDD_HHMMSSV_pairs.nc**. Commonly available NetCDF utilities such as ncdump or ncview may be used to view the contents of the output file. +If requested using the **nc_pairs_flag** dictionary in the configuration file, a NetCDF file containing the matched pair and forecast minus observation difference fields for each combination of variable type/level and masking region applied will be generated. The contents of this file are determined by the contents of the nc_pairs_flag dictionary. The output NetCDF file is named similarly to the other output files: **grid_stat_PREFIX_HHMMSSL_YYYYMMDD_HHMMSSV_pairs.nc**. Commonly available NetCDF utilities such as ncdump or ncview may be used to view the contents of the output file. The output NetCDF file contains the dimensions and variables shown in :numref:`table_GS_Dimensions_NetCDF_matched_pair_out` and :numref:`table_GS_var_NetCDF_matched_pair_out`. @@ -969,13 +969,13 @@ The output NetCDF file contains the dimensions and variables shown in :numref:`t * - NetCDF Dimension - Description * - Lat - - Dimension of the latitude (i.e. Number of grid points in the North-South direction) + - Dimension of the latitude (i.e., Number of grid points in the North-South direction) * - Lon - - Dimension of the longitude (i.e. Number of grid points in the East-West direction) + - Dimension of the longitude (i.e., Number of grid points in the East-West direction) .. role:: raw-html(raw) - :format: html + :format: html .. _table_GS_var_NetCDF_matched_pair_out: diff --git a/docs/Users_Guide/gsi-tools.rst b/docs/Users_Guide/gsi-tools.rst index 010166ab3e..a5e7b620cd 100644 --- a/docs/Users_Guide/gsi-tools.rst +++ b/docs/Users_Guide/gsi-tools.rst @@ -40,7 +40,7 @@ gsid2mpr has one required argument and accepts several optional ones. Required Arguments for gsid2mpr ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ -1. The **gsi_file_1 [gsi_file2 ... gsi_file_n]** argument indicates the GSI diagnostic files (conventional or radiance) to be reformatted. +1. The **gsi_file_1 [gsi_file_2 ... gsi_file_n]** argument indicates the GSI diagnostic files (conventional or radiance) to be reformatted. Optional Arguments for gsid2mpr ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ @@ -69,12 +69,12 @@ An example of the gsid2mpr calling sequence is shown below: -set_hdr MODEL GSI_MEM001 \ -outdir out -In this example, the GSID2MPR tool will process a single input file named **diag_conv_ges.mem001** file, set the output **MODEL** header column to **GSI_MEM001**, and write output to the **out** directory. The output file is named the same as the input file but a **.stat** suffix is added to indicate its format. +In this example, the GSID2MPR tool will process a single input file named **diag_conv_ges.mem001**, set the output **MODEL** header column to **GSI_MEM001**, and write output to the **out** directory. The output file is named the same as the input file but a **.stat** suffix is added to indicate its format. gsid2mpr Output --------------- -The GSID2MPR tool performs a simple reformatting step and thus requires no configuration file. It can read both conventional and radiance binary GSI diagnostic files. Support for additional GSI diagnostic file type may be added in future releases. Conventional files are determined by the presence of the string **conv** in the filename. Files that are not conventional are assumed to contain radiance data. Multiple files of either type may be passed in a single call to the GSID2MPR tool. For each input file, an output file will be generated containing the corresponding matched pair data. +The GSID2MPR tool performs a simple reformatting step and thus requires no configuration file. It can read both conventional and radiance binary GSI diagnostic files. Support for additional GSI diagnostic file types may be added in future releases. Conventional files are determined by the presence of the string **conv** in the filename. Files that are not conventional are assumed to contain radiance data. Multiple files of either type may be passed in a single call to the GSID2MPR tool. For each input file, an output file will be generated containing the corresponding matched pair data. The GSID2MPR tool writes the same set of MPR output columns for the conventional and radiance data types. However, it also writes additional columns at the end of the MPR line which depend on the input file type. Those additional columns are described in the following tables. @@ -126,7 +126,7 @@ The GSID2MPR tool writes the same set of MPR output columns for the conventional .. role:: raw-html(raw) - :format: html + :format: html .. list-table:: Format information for GSI Diagnostic Radiance MPR (Matched Pair) output line type. :widths: auto @@ -230,7 +230,7 @@ The GSID2MPR tool writes the same set of MPR output columns for the conventional - Double * - 60 - CTOP_PRS TC_PWAT - - Cloud top pressure (hPa) :raw-html:`
` Total column precip. water (km/m**2) (microwave only) + - Cloud top pressure (hPa) :raw-html:`
` Total column precip. water (kg/m**2) (microwave only) - Double * - 61 - TFND @@ -266,7 +266,7 @@ The GSID2MPR tool writes the same set of MPR output columns for the conventional - Double * - 69 - PRS_MAX_WGT - - Pressure of the maximum weighing function + - Pressure of the maximum weighting function - Double The gsid2mpr output may be passed to the Stat-Analysis tool to derive additional statistics. In particular, users should consider running the **aggregate_stat** job type to read MPR lines and compute partial sums (SL1L2), continuous statistics (CNT), contingency table counts (CTC), or contingency table statistics (CTS). Stat-Analysis has been enhanced to parse any extra columns found at the end of the input lines. Users can filter the values in those extra columns using the **-column_thresh**, **-column_str**, and **-column_str_exc** job command options. @@ -347,7 +347,7 @@ gsidens2orank Output The GSIDENS2ORANK tool performs a simple reformatting step and thus requires no configuration file. The multiple files passed to it are interpreted as members of the same ensemble. Therefore, each call to the tool processes exactly one ensemble. All input ensemble GSI diagnostic files must be of the same type. Mixing conventional and radiance files together will result in a runtime error. The GSIDENS2ORANK tool processes each ensemble member and keeps track of the observations it encounters. It constructs a list of the ensemble values corresponding to each observation and writes an output ORANK line listing the observation value, its rank, and all the ensemble values. The random number generator is used by the GSIDENS2ORANK tool to randomly assign a rank value in the case of ties. -The GSID2MPR tool writes the same set of ORANK output columns for the conventional and radiance data types. However, it also writes additional columns at the end of the ORANK line which depend on the input file type. The extra columns are limited to quantities which remain constant over all the ensemble members and are therefore largely a subset of the extra columns written by the GSID2MPR tool. Those additional columns are described in the following tables. +The GSIDENS2ORANK tool writes the same set of ORANK output columns for the conventional and radiance data types. However, it also writes additional columns at the end of the ORANK line which depend on the input file type. The extra columns are limited to quantities which remain constant over all the ensemble members and are therefore largely a subset of the extra columns written by the GSID2MPR tool. Those additional columns are described in the following tables. .. list-table:: Format information for GSI Diagnostic Conventional ORANK (Observation Rank) output line type. :widths: auto @@ -488,7 +488,7 @@ The GSID2MPR tool writes the same set of ORANK output columns for the convention - d(Tz)/d(Tr) - Double -The gsidens2orank output may be passed to the Stat-Analysis tool to derive additional statistics. In particular, users should consider running the **aggregate_stat** job type to read ORANK lines and ranked histograms (RHIST), probability integral transform histograms (PHIST), and spread-skill variance output (SSVAR). Stat-Analysis has been enhanced to parse any extra columns found at the end of the input lines. Users can filter the values in those extra columns using the **-column_thresh**, **-column_str**, and **-column_str_exc** job command options. +The gsidens2orank output may be passed to the Stat-Analysis tool to derive additional statistics. In particular, users should consider running the **aggregate_stat** job type to read ORANK lines and compute ranked histograms (RHIST), probability integral transform histograms (PHIST), and spread-skill variance output (SSVAR). Stat-Analysis has been enhanced to parse any extra columns found at the end of the input lines. Users can filter the values in those extra columns using the **-column_thresh**, **-column_str**, and **-column_str_exc** job command options. An example of the Stat-Analysis calling sequence is shown below: @@ -498,4 +498,4 @@ An example of the Stat-Analysis calling sequence is shown below: -job aggregate_stat -line_type ORANK -out_line_type RHIST \ -by fcst_var -column_thresh N_USE eq20 -In this example, the Stat-Analysis tool will read ORANK lines from **diag_conv_ges_ens_mean_orank.txt**, retain only those lines where the **N_USE** column indicates that all 20 ensemble members were used, and write ranked histogram (RHIST) output lines for each unique value of encountered in the **FCST_VAR** column. +In this example, the Stat-Analysis tool will read ORANK lines from **diag_conv_ges_ens_mean_orank.txt**, retain only those lines where the **N_USE** column indicates that all 20 ensemble members were used, and write ranked histogram (RHIST) output lines for each unique value encountered in the **FCST_VAR** column. diff --git a/docs/Users_Guide/index.rst b/docs/Users_Guide/index.rst index 9458c7a9bb..36946d9da2 100644 --- a/docs/Users_Guide/index.rst +++ b/docs/Users_Guide/index.rst @@ -4,7 +4,7 @@ User's Guide **Foreword: A note to MET users** -This User's guide is provided as an aid to users of the Model Evaluation Tools (MET). MET is a set of verification tools developed by the Developmental Testbed Center (DTC) for use by the numerical weather prediction community to help them assess and evaluate the performance of numerical weather predictions. It is also the core component of the unified METplus verification framework. More details about METplus can be found on the `METplus website `_. +This User's Guide is provided as an aid to users of the Model Evaluation Tools (MET). MET is a set of verification tools developed by the Developmental Testbed Center (DTC) for use by the numerical weather prediction community to help them assess and evaluate the performance of numerical weather predictions. It is also the core component of the unified METplus verification framework. More details about METplus can be found on the `METplus website `_. It is important to note here that MET is an evolving software package. This documentation describes the |release| release dated |release_date|. Previous releases of MET have occurred each year since 2008. Intermediate releases may include bug fixes. MET is also able to accept new modules contributed by the community. If you have code you would like to contribute, we will gladly consider your contribution. Please create a post in the `METplus GitHub Discussions Forum `_. We will then determine the maturity of the new verification method and coordinate the inclusion of the new module in a future version. @@ -35,61 +35,61 @@ Developmental Testbed Center. Available at: https://github.com/dtcenter/MET/rele **Acknowledgments** -We thank the National Science Foundation (NSF) along with three organizations within the National Oceanic and Atmospheric Administration (NOAA): 1) Office of Atmospheric Research (OAR); 2) Next Generation Global Prediction System project (NGGPS); and 3) United State Weather Research Program (USWRP), the United States Air Force (USAF), and the United States Department of Energy (DOE) for their support of this work. Funding for the development of MET-TC is from the NOAA's Hurricane Forecast Improvement Project (HFIP) through the Developmental Testbed Center (DTC). Funding for the expansion of capability to address many methods pertinent to global and climate simulations was provided by NOAA's Next Generation Global Prediction System (NGGPS) and NSF Earth System Model 2 (EaSM2) projects. We would like to thank James Franklin at the National Hurricane Center (NHC) for his insight into the original development of the existing NHC verification software. Thanks also go to the staff at the Developmental Testbed Center for their help, advice, and many types of support. We released METv1.0 in January 2008 and would not have made a decade of cutting-edge verification support without those who participated in the original MET planning workshops and the now dis-banded verification advisory group (Mike Baldwin, Matthew Sittel, Elizabeth Ebert, Geoff DiMego, Chris Davis, and Jason Knievel). +We thank the National Science Foundation (NSF) along with three organizations within the National Oceanic and Atmospheric Administration (NOAA): 1) Office of Atmospheric Research (OAR); 2) Next Generation Global Prediction System project (NGGPS); and 3) United States Weather Research Program (USWRP), the United States Air Force (USAF), and the United States Department of Energy (DOE) for their support of this work. Funding for the development of MET-TC is from the NOAA's Hurricane Forecast Improvement Project (HFIP) through the Developmental Testbed Center (DTC). Funding for the expansion of capability to address many methods pertinent to global and climate simulations was provided by NOAA's Next Generation Global Prediction System (NGGPS) and NSF Earth System Model 2 (EaSM2) projects. We would like to thank James Franklin at the National Hurricane Center (NHC) for his insight into the original development of the existing NHC verification software. Thanks also go to the staff at the Developmental Testbed Center for their help, advice, and many types of support. We released METv1.0 in January 2008 and would not have made a decade of cutting-edge verification support without those who participated in the original MET planning workshops and the now disbanded verification advisory group (Mike Baldwin, Matthew Sittel, Elizabeth Ebert, Geoff DiMego, Chris Davis, and Jason Knievel). The National Center for Atmospheric Research (NCAR) is sponsored by NSF. The DTC is sponsored by the National Oceanic and Atmospheric Administration (NOAA), the United States Air Force, and the National Science Foundation (NSF). NCAR is sponsored by the National Science Foundation (NSF). .. toctree:: - :titlesonly: - :numbered: 4 - - overview - release-notes - installation - data_io - config_options - config_options_tc - reformat_point - reformat_grid - gen-ens-prod - masking - point-stat - pair-stat - grid-stat - ensemble-stat - wavelet-stat - gsi-tools - stat-analysis - series-analysis - grid-diag - mode - mode-analysis - mode-td - met-tc_overview - tc-dland - tc-pairs - tc-diag - tc-stat - tc-gen - tc-rmw - rmw-analysis - plotting - refs - appendixA - appendixB - appendixC - appendixD - appendixE - appendixF - appendixG - appendixH + :titlesonly: + :numbered: 4 + + overview + release-notes + installation + data_io + config_options + config_options_tc + reformat_point + reformat_grid + gen-ens-prod + masking + point-stat + pair-stat + grid-stat + ensemble-stat + wavelet-stat + gsi-tools + stat-analysis + series-analysis + grid-diag + mode + mode-analysis + mode-td + met-tc_overview + tc-dland + tc-pairs + tc-diag + tc-stat + tc-gen + tc-rmw + rmw-analysis + plotting + refs + appendixA + appendixB + appendixC + appendixD + appendixE + appendixF + appendixG + appendixH .. only:: html - Indices and tables - ================== + Indices and tables + ================== - * :ref:`genindex` - * :ref:`search` + * :ref:`genindex` + * :ref:`search` diff --git a/docs/Users_Guide/installation.rst b/docs/Users_Guide/installation.rst index 8e6b54a1e6..5dd8a018c8 100644 --- a/docs/Users_Guide/installation.rst +++ b/docs/Users_Guide/installation.rst @@ -70,7 +70,7 @@ Users can take advantage of the compilation script to download and install all o libraries automatically, both required and conditionally required :ref:`compile_script_install`. -.. _suggested_external_utiliites: +.. _suggested_external_utilities: Suggested External Utilities ============================ @@ -85,7 +85,7 @@ They are not required for MET to function, but depending on the user’s intende * `Integrated Data Viewer (IDV) `_ for displaying gridded data, including GRIB and NetCDF * `ncview utility `_ - for viewing gridded NetCDF data (e.g. the output of pcp_combine) + for viewing gridded NetCDF data (e.g., the output of pcp_combine) .. _compile_script_install: @@ -106,7 +106,7 @@ format from GitHub, which the script will then install. To begin, create and change to a directory where the latest version of MET will be installed. Assuming that the following guidance uses “/d1” as the parent directory, a suggested format is a path to a “met” directory, followed by the version number -subdirectory (e.g. /d1/met/13.0.0). +subdirectory (e.g., /d1/met/13.0.0). Next, download the `compile_MET_all.sh `_ script and @@ -140,7 +140,7 @@ Now change directories to the one that was created from expanding the tar files: cd tar_files The next step will be to identify and download the latest MET release as a -tar file (e.g. v13.0.0.tar.gz) and place it in +tar file (e.g., v13.0.0.tar.gz) and place it in the *tar_files* directory. The file is available from the MET line under the “RECOMMENDED - COMPONENTS” section on the `METplus website `_ or @@ -175,134 +175,134 @@ directory. .. note:: Starting with MET-12.0.0, C++17 is the default C++ standard for MET due to the requirements of its dependent libraries. However, MET itself only makes use of C++11 features. - The ATLAS library (conditionally required for MET, if support for - unstructured grids is desired) - `versions 0.33.0 `_ - and later requires compiler support for the C++17 standard. + The ATLAS library (conditionally required for MET, if support for + unstructured grids is desired) + `versions 0.33.0 `_ + and later requires compiler support for the C++17 standard. - At this time, users with systems that do not yet support the C++17 - standard, can still compile MET with an older C++ standard, using an - older version of ATLAS, by adding the MET_CXX_STANDARD variable to - the environment configuration file as described in the **OPTIONAL** - section below. + At this time, users with systems that do not yet support the C++17 + standard, can still compile MET with an older C++ standard, using an + older version of ATLAS, by adding the MET_CXX_STANDARD variable to + the environment configuration file as described in the **OPTIONAL** + section below. Environment Variable Descriptions --------------------------------- .. dropdown:: REQUIRED - **TEST_BASE** – Format is */d1/met/13.0.0*. This is the MET - installation directory that was created - the beginning of, :numref:`compile_script_install` and contains the - **compile_MET_all.sh** script, **tar_files.tgz**, - and the *tar_files* directory from the untar command. - - **COMPILER** – Format is *compiler_version* (e.g. gnu_8.3.0). For the GNU family of compilers, - use “gnu”; for the Intel family of compilers, use “intel”, "intel-classic", - “intel-oneapi”, “ics”, “ips”, or “PrgEnv-intel”, - depending on the system. If using an Intel compiler, users that have also - set the **USE_MODULES** environment variable to TRUE should review the additional - information below for proper configuration file setup. In the past, support was - provided for the PGI family of compilers through “pgi”. However, this compiler - option is no longer actively tested. - - **MET_SUBDIR** – Format is */d1/met/13.0.0*. This is the location where the top-level MET - subdirectory will - be installed and is often set equivalent to **TEST_BASE** (e.g. ${TEST_BASE}). - - **MET_TARBALL** – Format is *v13.0.0.tar.gz*. This is the name of the downloaded MET tarball. - - **USE_MODULES** – Format is *TRUE* or *FALSE*. Set to FALSE if using a machine that does not use - modulefiles; set to TRUE if using a machine that does use modulefiles. For more information on - modulefiles, visit the `Wikipedia page `_. - If the **USE_MODULES** setting is set to true and the compiler is an Intel compiler, please - review the additional information below for proper configuration file setup. - - **PYTHON_MODULE** - Format is *PythonModuleName_version* (e.g. python_3.10.4). This environment variable - is only required if **USE_MODULES** = TRUE. To set properly, list the Python module to load - followed by an underscore and version number. For example, setting - **PYTHON_MODULE** =python_3.10.4 - will cause the script to run "module load python/3.10.4". + **TEST_BASE** – Format is */d1/met/13.0.0*. This is the MET + installation directory that was created + at the beginning of :numref:`compile_script_install` and contains the + **compile_MET_all.sh** script, **tar_files.tgz**, + and the *tar_files* directory from the untar command. + + **COMPILER** – Format is *compiler_version* (e.g., gnu_8.3.0). For the GNU family of compilers, + use “gnu”; for the Intel family of compilers, use “intel”, "intel-classic", + “intel-oneapi”, “ics”, “ips”, or “PrgEnv-intel”, + depending on the system. If using an Intel compiler, users that have also + set the **USE_MODULES** environment variable to TRUE should review the additional + information below for proper configuration file setup. In the past, support was + provided for the PGI family of compilers through “pgi”. However, this compiler + option is no longer actively tested. + + **MET_SUBDIR** – Format is */d1/met/13.0.0*. This is the location where the top-level MET + subdirectory will + be installed and is often set equivalent to **TEST_BASE** (e.g., ${TEST_BASE}). + + **MET_TARBALL** – Format is *v13.0.0.tar.gz*. This is the name of the downloaded MET tarball. + + **USE_MODULES** – Format is *TRUE* or *FALSE*. Set to FALSE if using a machine that does not use + modulefiles; set to TRUE if using a machine that does use modulefiles. For more information on + modulefiles, visit the `Wikipedia page `_. + If the **USE_MODULES** setting is set to true and the compiler is an Intel compiler, please + review the additional information below for proper configuration file setup. + + **PYTHON_MODULE** - Format is *PythonModuleName_version* (e.g., python_3.10.4). This environment variable + is only required if **USE_MODULES** = TRUE. To set properly, list the Python module to load + followed by an underscore and version number. For example, setting + **PYTHON_MODULE** =python_3.10.4 + will cause the script to run "module load python/3.10.4". .. dropdown:: ADDITIONAL SETTINGS FOR INTEL COMPILER USERS WITH THE USE_MODULES SETTING - It is necessary for the user to specify (in the install_met_env. config file) the - following environment variables if using the Intel compilers: - - | For non-oneAPI Intel compilers: - | - | export FC=ifort - | export F77=ifort - | export F90=ifort - | export CC=icc - | export CXX=icpc - - - | For oneAPI Intel compilers: - | - | export FC=ifx - | export F77=ifx - | export F90=ifx - | export CC=icx - | export CXX=icpx - - This is due to the machines allowing users to load a module but not setting these environment - variables as expected, leading to failed installations. For user convenience, additional - generic configuration files have been created that include these settings. Users with a - classic Intel compiler are encouraged to use the install_met_env.generic_intel_classic - configuration file, and users with a oneAPI Intel compiler should use the - install_met_env.generic_intel_oneapi configuration file. + It is necessary for the user to specify (in the install_met_env. config file) the + following environment variables if using the Intel compilers: + + | For non-oneAPI Intel compilers: + | + | export FC=ifort + | export F77=ifort + | export F90=ifort + | export CC=icc + | export CXX=icpc + + + | For oneAPI Intel compilers: + | + | export FC=ifx + | export F77=ifx + | export F90=ifx + | export CC=icx + | export CXX=icpx + + This is due to the machines allowing users to load a module but not setting these environment + variables as expected, leading to failed installations. For user convenience, additional + generic configuration files have been created that include these settings. Users with a + classic Intel compiler are encouraged to use the install_met_env.generic_intel_classic + configuration file, and users with a oneAPI Intel compiler should use the + install_met_env.generic_intel_oneapi configuration file. .. dropdown:: REQUIRED, IF COMPILING PYTHON EMBEDDING - **MET_PYTHON** – Format is */usr/local/python3*. - This is the location - containing the bin, include, lib, and share directories for Python. - - **MET_PYTHON_CC** - Format is -I followed by the directory containing - the Python include files (e.g. -I/usr/local/python3/include/python3.10). - This information may be obtained by - running :code:`python3-config --cflags`; - however, this command can, on certain systems, - provide too much information. - - **MET_PYTHON_LD** - Format is -L followed by the directory containing - the Python library - files then a space, then -l followed by the necessary Python - libraries to link to - (e.g. -L/usr/local/python3/lib/\\ -lpython3.10\\ - -lpthread\\ -ldl\\ -lutil\\ -lm). - The backslashes are necessary in the example shown because of - the spaces, which will be - recognized as the end of the value unless preceded by the “\\” - character. Alternatively, - a user can provide the value in quotations - (e.g. export MET_PYTHON_LD="-L/usr/local/python3/lib/ - -lpython3.10 -lpthread -ldl -lutil -lm"). - This information may be obtained by running - :code:`python3-config --ldflags --embed`; however, - this command can, on certain systems, provide too much information. + **MET_PYTHON** – Format is */usr/local/python3*. + This is the location + containing the bin, include, lib, and share directories for Python. + + **MET_PYTHON_CC** - Format is -I followed by the directory containing + the Python include files (e.g., -I/usr/local/python3/include/python3.10). + This information may be obtained by + running :code:`python3-config --cflags`; + however, this command can, on certain systems, + provide too much information. + + **MET_PYTHON_LD** - Format is -L followed by the directory containing + the Python library + files then a space, then -l followed by the necessary Python + libraries to link to + (e.g., -L/usr/local/python3/lib/\\ -lpython3.10\\ + -lpthread\\ -ldl\\ -lutil\\ -lm). + The backslashes are necessary in the example shown because of + the spaces, which will be + recognized as the end of the value unless preceded by the “\\” + character. Alternatively, + a user can provide the value in quotations + (e.g., export MET_PYTHON_LD="-L/usr/local/python3/lib/ + -lpython3.10 -lpthread -ldl -lutil -lm"). + This information may be obtained by running + :code:`python3-config --ldflags --embed`; however, + this command can, on certain systems, provide too much information. .. dropdown:: OPTIONAL - **export MAKE_ARGS="-j #"** – If there is a need to install external - libraries, or to attempt - to speed up the MET compilation process, this environmental - setting can be added to the - environment configuration file. Replace the # with the number - of cores to use - (as an integer) or simply specify "export MAKE_ARGS=-j" - with no integer argument to - start as many processes in parallel as possible. Note that Docker - has trouble compiling - without a specified value of cores to use. The automated MET - testing scripts in the - Docker environment have been successful with a value of - 5 (e.g. export MAKE_ARGS=”-j 5”). - - **export MET_CXX_STANDARD** - Specify the version of the supported - C++ standard. Values may be 11, 14, or 17. The default value is 17. - (e.g. export MET_CXX_STANDARD=11) + **export MAKE_ARGS="-j #"** – If there is a need to install external + libraries, or to attempt + to speed up the MET compilation process, this environmental + setting can be added to the + environment configuration file. Replace the # with the number + of cores to use + (as an integer) or simply specify "export MAKE_ARGS=-j" + with no integer argument to + start as many processes in parallel as possible. Note that Docker + has trouble compiling + without a specified value of cores to use. The automated MET + testing scripts in the + Docker environment have been successful with a value of + 5 (e.g., export MAKE_ARGS=”-j 5”). + + **export MET_CXX_STANDARD** - Specify the version of the supported + C++ standard. Values may be 11, 14, or 17. The default value is 17. + (e.g., export MET_CXX_STANDARD=11) External Library Handling in compile_MET_all.sh @@ -310,94 +310,94 @@ External Library Handling in compile_MET_all.sh .. dropdown:: IF THE USER WANTS TO HAVE THE COMPILATION SCRIPT COMPILE THE LIBRARY DEPENDENCIES - The **compile_MET_all.sh** script will compile and install MET and its - :ref:`required_external_libraries_to_build_MET`, if needed. - Note that if these libraries are already installed somewhere on the system, - MET will call and use the libraries that were installed by the script. + The **compile_MET_all.sh** script will compile and install MET and its + :ref:`required_external_libraries_to_build_MET`, if needed. + Note that if these libraries are already installed somewhere on the system, + MET will call and use the libraries that were installed by the script. .. dropdown:: IF THE USER ALREADY HAS THE LIBRARY DEPENDENCIES INSTALLED - If the required external library dependencies have already been installed and don’t - need to be reinstalled, or if compiling MET on a machine that uses modulefiles and - the user would like to make use of the existing dependent libraries on that machine, - there are more environment variables that need to be set to let MET know where those - library and header files are. The following environment variables need to be added - to the environment configuration file: - - +-------------------+--------------------------------+------------------------------+ - | **Feature** | **Configuration Option** | **Environment Variables** | - +===================+================================+==============================+ - | *Always* | | MET_BUFRLIB, | - | | | | - | *Required* | | BUFRLIB_NAME, | - | | | | - | | | MET_PROJ, | - | | | | - | | | MET_HDF5, | - | | | | - | | | MET_NETCDF, | - | | | | - | | | MET_GSL | - +-------------------+--------------------------------+------------------------------+ - | *Optional* | :code:`--enable-all` or | MET_GRIB2CLIB, | - | | | | - | GRIB2 | :code:`--enable-grib2` | MET_GRIB2CINC, | - | | | | - | Support | | GRIB2CLIB_NAME, | - | | | | - | | | LIB_JASPER, | - | | | | - | | | LIB_PNG, | - | | | | - | | | LIB_AEC, | - | | | | - | | | LIB_Z | - +-------------------+--------------------------------+------------------------------+ - | *Optional* | :code:`--enable-all` or | MET_PYTHON_BIN_EXE, | - | | | | - | Python | :code:`--enable-python` | MET_PYTHON_CC, | - | | | | - | Support | | MET_PYTHON_LD | - +-------------------+--------------------------------+------------------------------+ - | *Optional* | :code:`--enable-all` or | MET_ATLAS, | - | | | | - | Unstructured Grid | :code:`--enable-ugrid` | MET_ECKIT | - | | | | - | Support | | | - +-------------------+--------------------------------+------------------------------+ - | *Optional* | :code:`--enable-all` or | MET_HDF | - | | | | - | LIDAR2NC | :code:`--enable-lidar2nc` | | - | | | | - | Support | | | - +-------------------+--------------------------------+------------------------------+ - | *Optional* | :code:`--enable-all` or | MET_HDF, | - | | | | - | MODIS | :code:`--enable-modis` | MET_HDFEOS | - | | | | - | Support | | | - +-------------------+--------------------------------+------------------------------+ - | *Optional* | :code:`--enable-profiler` | | - | | | | - | Profiler | | | - | | | | - | Support | | | - +-------------------+--------------------------------+------------------------------+ - - Generally speaking, for each library there is a set of three - environment variables that can - describe the locations: - **$MET_**, **$MET_INC** and **$MET_LIB**. - - The $MET_ environment variable can be used if the external library is - installed such that there is a main directory which has a subdirectory called - *lib* containing the library files and another subdirectory called *include* - containing the include files. - - Alternatively, the $MET_INC and $MET_LIB environment variables are used if the - library and include files for an external library are installed in separate locations. - In this case, both environment variables must be specified and the associated - $MET_ variable will be ignored. + If the required external library dependencies have already been installed and don’t + need to be reinstalled, or if compiling MET on a machine that uses modulefiles and + the user would like to make use of the existing dependent libraries on that machine, + there are more environment variables that need to be set to let MET know where those + library and header files are. The following environment variables need to be added + to the environment configuration file: + + +-------------------+--------------------------------+------------------------------+ + | **Feature** | **Configuration Option** | **Environment Variables** | + +===================+================================+==============================+ + | *Always* | | MET_BUFRLIB, | + | | | | + | *Required* | | BUFRLIB_NAME, | + | | | | + | | | MET_PROJ, | + | | | | + | | | MET_HDF5, | + | | | | + | | | MET_NETCDF, | + | | | | + | | | MET_GSL | + +-------------------+--------------------------------+------------------------------+ + | *Optional* | :code:`--enable-all` or | MET_GRIB2CLIB, | + | | | | + | GRIB2 | :code:`--enable-grib2` | MET_GRIB2CINC, | + | | | | + | Support | | GRIB2CLIB_NAME, | + | | | | + | | | LIB_JASPER, | + | | | | + | | | LIB_PNG, | + | | | | + | | | LIB_AEC, | + | | | | + | | | LIB_Z | + +-------------------+--------------------------------+------------------------------+ + | *Optional* | :code:`--enable-all` or | MET_PYTHON_BIN_EXE, | + | | | | + | Python | :code:`--enable-python` | MET_PYTHON_CC, | + | | | | + | Support | | MET_PYTHON_LD | + +-------------------+--------------------------------+------------------------------+ + | *Optional* | :code:`--enable-all` or | MET_ATLAS, | + | | | | + | Unstructured Grid | :code:`--enable-ugrid` | MET_ECKIT | + | | | | + | Support | | | + +-------------------+--------------------------------+------------------------------+ + | *Optional* | :code:`--enable-all` or | MET_HDF | + | | | | + | LIDAR2NC | :code:`--enable-lidar2nc` | | + | | | | + | Support | | | + +-------------------+--------------------------------+------------------------------+ + | *Optional* | :code:`--enable-all` or | MET_HDF, | + | | | | + | MODIS | :code:`--enable-modis` | MET_HDFEOS | + | | | | + | Support | | | + +-------------------+--------------------------------+------------------------------+ + | *Optional* | :code:`--enable-profiler` | | + | | | | + | Profiler | | | + | | | | + | Support | | | + +-------------------+--------------------------------+------------------------------+ + + Generally speaking, for each library there is a set of three + environment variables that can + describe the locations: + **$MET_**, **$MET_INC** and **$MET_LIB**. + + The $MET_ environment variable can be used if the external library is + installed such that there is a main directory which has a subdirectory called + *lib* containing the library files and another subdirectory called *include* + containing the include files. + + Alternatively, the $MET_INC and $MET_LIB environment variables are used if the + library and include files for an external library are installed in separate locations. + In this case, both environment variables must be specified and the associated + $MET_ variable will be ignored. Executing the compile_MET_all.sh script --------------------------------------- @@ -441,7 +441,7 @@ To confirm that MET was installed successfully, run the following command from t If no errors are returned, the installation was successful. Due to the highly variable nature of hardware systems, users may encounter issues during -the installation process that result in MET not being installed. If this occurs please +the installation process that result in MET not being installed. If this occurs, please first recheck that the location of all the necessary data files and scripts is correct. Next, recheck the environment variables in the environment configuration file and ensure there are no spelling errors or improperly set variables. @@ -468,10 +468,10 @@ down system environment settings and meet with success faster) alike. MET has numerous version images for Docker users and continues to be released as images at the same interval as system releases. While the advantages of Docker can -make it an appealing installation route for first time users, it does require +make it an appealing installation route for first-time users, it does require privileged user access that will result in an unsuccessful installation if not available. Please ensure the user has high system access -(e.g. admin access) before attempting this method. +(e.g., admin access) before attempting this method. Installing Docker ----------------- @@ -523,7 +523,7 @@ the same way the latest image of MET was pulled: docker run -it --rm dtcenter/met:13.0.0 /bin/bash -If the usage MET via Docker images was successful, it is highly +If the usage of MET via Docker images was successful, it is highly recommended to move on to using the METplus wrappers of the tools, which have their own Docker image. @@ -564,7 +564,7 @@ Loading the Latest MET Image Similar to Docker, Apptainer will build the container based off of the MET image in a single command. To accomplish this, Apptainer’s “Swiss army knife” :code:`build` -command is used. Use the the latest MET version number in +command is used. Use the latest MET version number in conjunction with :code:`build` to make the container: diff --git a/docs/Users_Guide/masking.rst b/docs/Users_Guide/masking.rst index 9d75704820..86638f7992 100644 --- a/docs/Users_Guide/masking.rst +++ b/docs/Users_Guide/masking.rst @@ -65,7 +65,7 @@ Required Arguments for gen_vx_mask .. note:: - While multiple **-type** mask types can be requested in a single run, all requested masking types must use the same **mask_file** setting. + While multiple **-type** mask types can be requested in a single run, all requested masking types must use the same **mask_file** setting. Optional Arguments for gen_vx_mask ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ @@ -116,7 +116,7 @@ The Gen-Vx-Mask tool supports the following types of masking region definition s 2. Polyline XY (**poly_xy**) masking reads an input ASCII file containing Lat/Lon locations. It converts the polyline Lat/Lon locations into grid X/Y space and connects the first and last points. It selects grid points whose X/Y location falls inside that polyline in X/Y space. This option is useful when defining geographic subregions of a domain. -3. Box (**box**) masking reads an input ASCII file containing Lat/Lon locations and draws a box around each point. The height and width of the box is specified by the **-height** and **-width** command line options in grid units. For a square, only one of **-height** or **-width** needs to be used. +3. Box (**box**) masking reads an input ASCII file containing Lat/Lon locations and draws a box around each point. The height and width of the box are specified by the **-height** and **-width** command line options in grid units. For a square, only one of **-height** or **-width** needs to be used. 4. Circle (**circle**) masking reads an input ASCII file containing Lat/Lon locations and for each grid point, computes the minimum great-circle arc distance in kilometers to those points. If the **-thresh** command line option is not used, the minimum distance value for each grid point will be written to the output. If it is used, only those grid points whose minimum distance meets the threshold criteria will be selected. This option is useful when defining areas within a certain radius of radar locations. @@ -124,21 +124,21 @@ The Gen-Vx-Mask tool supports the following types of masking region definition s 6. Grid (**grid**) masking reads an input gridded data file, extracts the field specified using its grid definition, and selects grid points falling inside that grid. This option is useful when using a model nest to define the corresponding area of the parent domain. -7. Data (**data**) masking reads an input gridded data file, extracts the field specified using the **-mask_field** command line option, thresholds the data using the **-thresh** command line option, and selects grid points which meet that threshold criteria. The option is useful when thresholding topography to define a mask based on elevation or when threshold land use to extract a particular category. +7. Data (**data**) masking reads an input gridded data file, extracts the field specified using the **-mask_field** command line option, thresholds the data using the **-thresh** command line option, and selects grid points which meet that threshold criteria. The option is useful when thresholding topography to define a mask based on elevation or when thresholding land use to extract a particular category. 8. Solar altitude (**solar_alt**) and solar azimuth (**solar_azi**) masking computes the solar altitude and azimuth values in degrees at each grid point for the time defined by the **mask_file** setting. **mask_file** may either be set to an explicit time string in YYYYMMDD[_HH[MMSS]] UTC format or to a gridded data file. If set to a gridded data file, the **-mask_field** command line option specifies the field of data whose valid time should be used. If the **-thresh** command line option is not used, the raw solar altitude or azimuth degrees for each grid point will be written to the output. If it is used, the resulting binary mask field will be written. This option is useful when defining a day/night mask. -9. Solar time (**solar_time**) masking computes the solar time in decimal hours at each grid point for the for the time defined by the **mask_file** setting, as described above. The solar hours of the day range from 0 to 24, with a value of 12 indicating solar noon. Note that solar time is based only on longitude. If the **-thresh** command line option is not used, the raw solar time hours will be written to the output. +9. Solar time (**solar_time**) masking computes the solar time in decimal hours at each grid point for the time defined by the **mask_file** setting, as described above. The solar hours of the day range from 0 to 24, with a value of 12 indicating solar noon. Note that solar time is based only on longitude. If the **-thresh** command line option is not used, the raw solar time hours will be written to the output. 10. Latitude (**lat**) and longitude (**lon**) masking computes the latitude and longitude value at each grid point. This logic only requires the definition of the grid, specified by the **input_file**. Technically, the **mask_file** is not needed, but a value must be specified for the command line to parse correctly. Users are advised to simply repeat the **input_file** setting twice. If the **-thresh** command line option is not used, the raw latitude or longitude values for each grid point will be written to the output. This option is useful when defining latitude or longitude bands over which to compute statistics. 11. Shapefile (**shape**) masking uses closed polygons taken from an ESRI shapefile to define the masking region. Gen-Vx-Mask reads the shapefile with the ".shp" suffix and extracts the latitude and longitudes of the vertices. The shapefile must consist of closed polygons rather than polylines, points, or any of the other data types that shapefiles support. When the **-shape_str** command line option is used, Gen-Vx-Mask also reads metadata from the corresponding dBASE file with the ".dbf" suffix. - Shapefiles usually contain more than one polygon, and the user must select which of these shapes should be used. The **-shapeno n** and **-shape_str name string** command line options enable the user to select one or more polygons from the shapefile. For **-shape n**, **n** is a comma-separated list of integer shape indices to be used. Note that these values are zero-based. So the first polygon in the shapefile is shape number 0, the second polygon in the shapefile is shape number 1, etc. For example, **-shapeno 0,1,2** uses the first three shapes in the shapefile. When multiple shapes are specified, the mask is defined as their union. So all grid points falling inside at least one of the specified shapes are included in the mask. + Shapefiles usually contain more than one polygon, and the user must select which of these shapes should be used. The **-shapeno n** and **-shape_str name string** command line options enable the user to select one or more polygons from the shapefile. For **-shapeno n**, **n** is a comma-separated list of integer shape indices to be used. Note that these values are zero-based. So the first polygon in the shapefile is shape number 0, the second polygon in the shapefile is shape number 1, etc. For example, **-shapeno 0,1,2** uses the first three shapes in the shapefile. When multiple shapes are specified, the mask is defined as their union. So all grid points falling inside at least one of the specified shapes are included in the mask. - For the user's convenience, some utilities that perform human-readable screen dumps of shapefile contents are provided with MET. The **gis_dump_shp**, **gis_dump_shx**, and **gis_dump_dbf** tools enable the user to examine the contents of these shapefiles. In particular, the **gis_dump_dbf** tool prints the name and values of the metadata for each record. The **-shape_str** command line option filters the shapes using the attributes listed in the **gis_dump_dbf** output, and requires two arguments. The **name** argument is set to any valid shapefile attribute, and the **string** argument is a comma-separated list of values to be matched. An example of using **-shape_str** is **-shape_str CONTINENT Europe**, which will match all "CONTINENT" attribues that have the string "Europe" in them. Strings that contain embedded whitespace should be enclosed in single quotes. Also note that case insensitive matching is used. For example, when using a global country outline shapefile, **-shape_str NAME 'united kingdom,united states of america'** matches the "NAME" attributes that have both "United Kingdom" and "United States of America" in them. If **-shape_str** is used multiple times, only shapes matching all the named attributes will be used. For example, **-shape_str CONTINENT Europe -shape_str NAME Spain,Portugal** will only match shapes where the "CONTINENT" attrinute contains "Europe "and the "NAME" attribute contains "Spain" or "Portugal". If a user wishes, they can combine both the **-shape_str** and **-shapeno** options. In this case, the union of all matches from the shapefile will be used. + For the user's convenience, some utilities that perform human-readable screen dumps of shapefile contents are provided with MET. The **gis_dump_shp**, **gis_dump_shx**, and **gis_dump_dbf** tools enable the user to examine the contents of these shapefiles. In particular, the **gis_dump_dbf** tool prints the name and values of the metadata for each record. The **-shape_str** command line option filters the shapes using the attributes listed in the **gis_dump_dbf** output, and requires two arguments. The **name** argument is set to any valid shapefile attribute, and the **string** argument is a comma-separated list of values to be matched. An example of using **-shape_str** is **-shape_str CONTINENT Europe**, which will match all "CONTINENT" attributes that have the string "Europe" in them. Strings that contain embedded whitespace should be enclosed in single quotes. Also note that case insensitive matching is used. For example, when using a global country outline shapefile, **-shape_str NAME 'united kingdom,united states of america'** matches the "NAME" attributes that have both "United Kingdom" and "United States of America" in them. If **-shape_str** is used multiple times, only shapes matching all the named attributes will be used. For example, **-shape_str CONTINENT Europe -shape_str NAME Spain,Portugal** will only match shapes where the "CONTINENT" attribute contains "Europe" and the "NAME" attribute contains "Spain" or "Portugal". If a user wishes, they can combine both the **-shape_str** and **-shapeno** options. In this case, the union of all matches from the shapefile will be used. -The polyline, polyline XY, box, circle, and track masking methods all read an ASCII file containing Lat/Lon locations. Those files must contain a string, which defines the name of the masking region, followed by a series of whitespace-separated latitude (degrees north) and longitude (degree east) values. +The polyline, polyline XY, box, circle, and track masking methods all read an ASCII file containing Lat/Lon locations. Those files must contain a string, which defines the name of the masking region, followed by a series of whitespace-separated latitude (degrees north) and longitude (degrees east) values. Logic for gen_vx_mask ^^^^^^^^^^^^^^^^^^^^^ @@ -189,11 +189,11 @@ An example of defining the northwest hemisphere of the earth, as defined by lati -intersection -name nw_hemisphere -The Gen-Vx-Mask tool to be run iteratively on its own output using different **mask_file** settings to generate complex masking areas. The **-union, -intersection**, and **-symdiff** options control the logic for combining the input field and current mask values at each grid point. For example, one could define a complex masking region by selecting grid points with an elevation greater than 1000 meters within a Contiguous United States geographic region by doing the following: +The Gen-Vx-Mask tool can be run iteratively on its own output using different **mask_file** settings to generate complex masking areas. The **-union, -intersection**, and **-symdiff** options control the logic for combining the input field and current mask values at each grid point. For example, one could define a complex masking region by selecting grid points with an elevation greater than 1000 meters within a Contiguous United States geographic region by doing the following: • Run Gen-Vx-Mask to apply data masking by thresholding a field of topography greater than 1000 meters. -• Run Gen-Vx-Mask a second time on the output from the first call and applying polyline masking to define the geographic area of interest. Use the **-intersection** option to only select grid points whose value is non-zero in both the input field and the current mask. +• Run Gen-Vx-Mask a second time on the output from the first call and apply polyline masking to define the geographic area of interest. Use the **-intersection** option to only select grid points whose value is non-zero in both the input field and the current mask. An example of this Gen-Vx-Mask calling sequence is shown below: @@ -214,4 +214,4 @@ Here, Gen-Vx-Mask uses the **data** masking type to read topography data (**TOPO Feature-Relative Methods ======================== -This section contains a description of several methods that may be used to perform feature-relative (or event -based) evaluation. The methodology pertains to examining the environment surrounding a particular feature or event such as a tropical, extra-tropical cyclone, convective cell, snow-band, etc. Several approaches are available for these types of investigations including applying masking described above (e.g. circle or box) or using the FORCE interpolation method in the regrid configuration option (see :numref:`config_options`). These methods generally require additional scripting, including potentially storm-track identification, outside of MET to be paired with the features of the MET tools. METplus may be used to execute this type of analysis. Please refer to the `METplus User's Guide `_. +This section contains a description of several methods that may be used to perform feature-relative (or event-based) evaluation. The methodology pertains to examining the environment surrounding a particular feature or event such as a tropical, extra-tropical cyclone, convective cell, snow-band, etc. Several approaches are available for these types of investigations including applying masking described above (e.g., circle or box) or using the FORCE interpolation method in the regrid configuration option (see :numref:`config_options`). These methods generally require additional scripting, including potentially storm-track identification, outside of MET to be paired with the features of the MET tools. METplus may be used to execute this type of analysis. Please refer to the `METplus User's Guide `_. diff --git a/docs/Users_Guide/met-tc_overview.rst b/docs/Users_Guide/met-tc_overview.rst index 635a14f802..d3915d9992 100644 --- a/docs/Users_Guide/met-tc_overview.rst +++ b/docs/Users_Guide/met-tc_overview.rst @@ -101,7 +101,7 @@ BASIN, CY, YYYYMMDDHH, TECHNUM/MIN, TECH, TAU, LatN/S, LonE/W, VMAX, MSLP, TY, R The TC-Pairs tool expects two input data sources in order to generate matched pairs and subsequent error statistics. The expected input for MET-TC is an ATCF format file from model output, or the operational aids files with the operational model output for the 'adeck' and the NHC best track analysis (BEST) for the 'bdeck'. The BEST is a subjectively smoothed representation of the storm's location and intensity over its lifetime. The track and intensity values are based on a retrospective assessment of all available observations of the storm. -The BEST is in ATCF file format and contains all the above listed common fields. Given the reference dataset is expected in ATCF file format, any second ATCF format file from model output or operational model output from the NHC aids files can be supplied as well. The expected use of the TC-Pairs tool is to generate matched pairs between model output and the BEST. Note that some of the columns in the TC-Pairs output are populated based on the BEST information (e.g. storm category), therefore use of a different baseline may reduce the available filtering options. +The BEST is in ATCF file format and contains all the above listed common fields. Given the reference dataset is expected in ATCF file format, any second ATCF format file from model output or operational model output from the NHC aids files can be supplied as well. The expected use of the TC-Pairs tool is to generate matched pairs between model output and the BEST. Note that some of the columns in the TC-Pairs output are populated based on the BEST information (e.g., storm category), therefore use of a different baseline may reduce the available filtering options. All operational model aids and the BEST can be obtained from the `NHC ftp server. `_ diff --git a/docs/Users_Guide/mode-analysis.rst b/docs/Users_Guide/mode-analysis.rst index 0bda0f0b85..b6b7c18672 100644 --- a/docs/Users_Guide/mode-analysis.rst +++ b/docs/Users_Guide/mode-analysis.rst @@ -14,7 +14,7 @@ Users may wish to summarize multiple ASCII files produced by MODE across many ca Scientific and Statistical Aspects ================================== -The MODE-Analysis tool operates in two modes, called "summary" and "bycase". In summary mode, the user specifies on the command line the MODE output columns of interest as well as filtering criteria that determine which input lines should be used. For example, a user may be interested in forecast object areas, but only if the object was matched, and only if the object centroid is inside a particular region. The summary statistics generated for each specified column of data are the minimum, maximum, mean, standard deviation, and the 10th, 25th, 50th, 75th and 90th percentiles. In addition, the user may specify a "dump'" file: the individual MODE lines used to produce the statistics will be written to this file. This option provides the user with a filtering capability. The dump file will consist only of lines that match the specified criteria. +The MODE-Analysis tool operates in two modes, called "summary" and "bycase". In summary mode, the user specifies on the command line the MODE output columns of interest as well as filtering criteria that determine which input lines should be used. For example, a user may be interested in forecast object areas, but only if the object was matched, and only if the object centroid is inside a particular region. The summary statistics generated for each specified column of data are the minimum, maximum, mean, standard deviation, and the 10th, 25th, 50th, 75th and 90th percentiles. In addition, the user may specify a "dump" file: the individual MODE lines used to produce the statistics will be written to this file. This option provides the user with a filtering capability. The dump file will consist only of lines that match the specified criteria. The other option for operating the analysis tool is "bycase". Given initial and final values for forecast lead time, the tool will output, for each valid time in the interval, the matched area, unmatched area, and the number of forecast and observed objects that were matched or unmatched. For the areas, the user can specify forecast or observed objects, and also simple or cluster objects. A dump file may also be specified in this mode. @@ -49,7 +49,7 @@ The MODE-Analysis tool has two required arguments and can accept several optiona Required Arguments for mode_analysis: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ -1. The **-lookin path** specifies the name of a specific STAT file (any file ending in .stat) or the name of a directory where the Stat-Analysis tool will search for STAT files. This option may be used multiple times to specify multiple locations. +1. The **-lookin path** specifies the name of a specific MODE object file (any file ending in _obj.txt) or the name of a directory where the MODE-Analysis tool will search for MODE object files. This option may be used multiple times to specify multiple locations. 2. The MODE-Analysis tool can perform two basic types of jobs **-summary** or **-bycase**. Exactly one of these job types must be specified. @@ -590,7 +590,7 @@ This option prints the usage message. mode_analysis Configuration File -------------------------------- -To use the MODE-Analysis tool, the user must un-comment the options in the configuration file to apply them and comment out unwanted options. The options in the configuration file for the MODE-Analysis tools are the same as the MODE command line options described in :numref:`mode_analysis-usage`. +To use the MODE-Analysis tool, the user must un-comment the options in the configuration file to apply them and comment out unwanted options. The options in the configuration file for the MODE-Analysis tool are the same as the MODE command line options described in :numref:`mode_analysis-usage`. The parameters that are set in the configuration file either add to or override parameters that are set on the command line. For the "set string" and "set integer type" options enclosed in brackets, the values specified in the configuration file are added to any values set on the command line. For the "toggle" and "min/max type" options, the values specified in the configuration file override those set on the command line. diff --git a/docs/Users_Guide/mode-td.rst b/docs/Users_Guide/mode-td.rst index 64e2475acf..4e89c37808 100644 --- a/docs/Users_Guide/mode-td.rst +++ b/docs/Users_Guide/mode-td.rst @@ -16,11 +16,11 @@ MODE Time Domain (MTD) is an extension of the MODE object-based approach to veri .. figure:: figure/mtd-3d_color.png - MTD Spacetime Objects + MTD Spacetime Objects A plot of some MTD precipitation objects is shown over the United States in :numref:`mtd-3d_color`. The colors indicate longitude, with red in the east moving through the spectrum to blue in the west. Time increases vertically in this plot (and in most of the spacetime diagrams in this users' guide). A few things are worthy of note in this figure. First, the tendency of storm systems to move from west to east over time shows up clearly. Second, tracking of storm objects over time is easily done: if we want to know if a storm at one time is a later version of a storm at an earlier time, we need only see if they are part of the same 3D spacetime object. Lastly, storms splitting up or merging over time are handled easily by this method. -The 2D (or traditional) MODE approach to object-base verification enabled users to analyze forecasts in terms of location errors, intensity errors and shape, size and orientation errors. MTD retains all of that capability, and adds new classes of forecast errors involving time information: speed and direction errors, buildup and decay errors, and timing and duration errors. This opens up new ways of analyzing forecast quality. +The 2D (or traditional) MODE approach to object-based verification enabled users to analyze forecasts in terms of location errors, intensity errors and shape, size and orientation errors. MTD retains all of that capability, and adds new classes of forecast errors involving time information: speed and direction errors, buildup and decay errors, and timing and duration errors. This opens up new ways of analyzing forecast quality. In the past, many MET users have performed separate MODE runs at a series of forecast valid times and analyzed the resulting object attributes, matches and merges as functions of time in an effort to incorporate temporal information in assessments of forecast quality. MTD was developed as a way to address this need in a more systematic way. Most of the information obtained from such multiple coordinated MODE runs can be obtained more simply from MTD. @@ -53,7 +53,7 @@ The most basic change is to use a square convolution filter rather than the circ .. figure:: figure/mtd-two_r_plus_one.png - Convolution Region + Convolution Region Another change is that we do not allow any bad data in the convolution square. In MODE, the user may specify what percentage of bad data in the convolution region is permissible, and it will rescale the value of the filter accordingly for each data point. For the sake of speed, MTD requires that there be no bad data in the convolution region. If any bad data exists in the region, the convolved value there is set to a bad data flag. @@ -70,7 +70,7 @@ The vector **velocity** :math:`(v_x, v_y)` is obtained by fitting a line to a 3D .. figure:: figure/mtd-velocity.png - Velocity + Velocity The spatial orientation of an object (what traditional MODE calls the **axis angle** of an object) is gotten by fitting a plane to an object. As with the case of velocity, our optimization criterion is that the sum of the squares of the spatial distances from each point of the object to the plane be minimized. @@ -80,13 +80,13 @@ The spatial orientation of an object (what traditional MODE calls the **axis ang .. figure:: figure/mtd-axis_3d.png - 3D axis + 3D axis -A simple integer count of the number of grid squares in an object for all of it's lifetime gives the **volume** of the object. Remember that while we're working in three dimensions, one of the dimensions is non-spatial, so one should not attempt to convert this to a volume in, e.g., :math:`\text{km}^3`. +A simple integer count of the number of grid squares in an object for all of its lifetime gives the **volume** of the object. Remember that while we're working in three dimensions, one of the dimensions is non-spatial, so one should not attempt to convert this to a volume in, e.g., :math:`\text{km}^3`. The **start time** and **end time** of an object are attributes as well. These are integers reflecting at which time step an object starts and ends. These values are zero-based, so for example, if an object comes into existence at the :math:`\text{3}^{rd}` time step and lasts until the :math:`\text{9}^{th}` time step, then the start time and end time will be listed as 2 and 8, respectively. Note that this object has a lifetime of 7 time steps, not 6. -**Centroid distance traveled** is the total great circle distance, in kilometers, traveled by the 2D spatial centroid over the lifetime of the object. In other words, at each time :math:`t` for which the 3D object exists, the set of points in the object also have that value of :math:`t` will together form a 2D spatial object. That 2D object will have a spatial centroid, which will move around as :math:`t` varies. This attribute represents this total 2D centroid movement over time. +**Centroid distance traveled** is the total great circle distance, in kilometers, traveled by the 2D spatial centroid over the lifetime of the object. In other words, at each time :math:`t` for which the 3D object exists, the set of points in the object that also have that value of :math:`t` will together form a 2D spatial object. That 2D object will have a spatial centroid, which will move around as :math:`t` varies. This attribute represents this total 2D centroid movement over time. Finally, MTD calculates several **intensity percentiles** of the raw data values inside each object. Not all of the attributes are purely geometrical. @@ -101,9 +101,9 @@ The **spatial centroid distance** is the purely spatial part of the centroid sep .. math:: \sqrt{(\bar{x_1} - \bar{x_2})^2 + (\bar{y_1} - \bar{y_2})^2 } -The **time centroid delta** is the difference between the time coordinates of the centroid. Since this is a simple difference, it can be either positive or negative. +The **time centroid delta** is the difference between the time coordinates of the centroid. Unlike the other deltas, it is computed as "observed minus forecast". Since this is a simple difference, it can be either positive or negative. -The **axis difference** is smaller of the two angles that the two spatial axis planes make with each other. :numref:`mtd-axis_diff` shows the idea. In the figure, the axis angle would be reported as angle :math:`\alpha`, not angle :math:`\beta`. +The **axis difference** is the smaller of the two angles that the two spatial axis planes make with each other. :numref:`mtd-axis_diff` shows the idea. In the figure, the axis angle would be reported as angle :math:`\alpha`, not angle :math:`\beta`. **Speed delta** and **direction difference** are obtained from the velocity vectors of the two objects. Speed delta is the difference in the lengths of the vectors, and direction difference is the angle that the two vectors make with each other. @@ -121,7 +121,7 @@ Finally, the **total interest** gives the result of the fuzzy-logic matching an .. figure:: figure/mtd-axis_diff.png - Axis Angle Difference + Axis Angle Difference 2D Constant-Time Attributes @@ -154,7 +154,7 @@ If we consider two distinct nodes in a graph to be related if there is a path co .. figure:: figure/mtd-basic_graph.png - Basic Graph Example + Basic Graph Example We have barely scratched the surface of the enormous subject of graph theory, but this will suffice for our purposes. How does MTD use graphs? Essentially the simple forecast and observed objects become nodes in a graph. Each pair of objects that have sufficiently high total interest (as determined by the fuzzy logic engine) generates an edge connecting the two corresponding nodes in the graph. The graph is then partitioned into equivalence classes using path connectivity (as explained above), and the resulting equivalence classes determine the matches and merges. @@ -174,7 +174,7 @@ To summarize: Any forecast simple objects that find themselves in the same equiv .. figure:: figure/mtd-2d_example.png - Match & Merge Example + Match & Merge Example Practical Information @@ -242,7 +242,7 @@ In this example, the MODE-TD tool will read in a list of forecast GRIB files in MTD Configuration File ---------------------- -The default configuration file for the MODE tool, **MODEConfig_default**, can be found in the installed *share/met/config* directory. Another version of the configuration file is provided in *scripts/config*. We encourage users to make a copy of the configuration files prior to modifying their contents.Most of the entries in the MTD configuration file should be familiar from the corresponding file for MODE. This initial beta release of MTD does not offer all the tunable options that MODE has accumulated over the years, however. In this section, we will not bother to repeat explanations of config file details that are exactly the same as those in MODE; we will only explain those elements that are different from MODE, and those that are unique to MTD. +The default configuration file for the MTD tool, **MTDConfig_default**, can be found in the installed *share/met/config* directory. Another version of the configuration file is provided in *scripts/config*. We encourage users to make a copy of the configuration files prior to modifying their contents. Most of the entries in the MTD configuration file should be familiar from the corresponding file for MODE. This initial beta release of MTD does not offer all the tunable options that MODE has accumulated over the years, however. In this section, we will not bother to repeat explanations of config file details that are exactly the same as those in MODE; we will only explain those elements that are different from MODE, and those that are unique to MTD. ______________________ @@ -275,7 +275,7 @@ ______________________ obs = fcst; total_interest_thresh = 0.7; -The configuration options listed above are common to many MODE and are described in :numref:`MODE-configuration-file`. +The configuration options listed above are common to MODE and are described in :numref:`MODE-configuration-file`. The **conv_time_window** entry is a dictionary defining how much smoothing in time should be done. The **beg** and **end** entries are integers defining how many time steps should be used before and after the current time. The default setting of **beg = -1; end = 1;** uses one time step before and after. Setting them both to 0 effectively disables smoothing in time. @@ -293,7 +293,7 @@ ______________________ min_volume = 2000; -The **min_volume** entry tells MTD to throw away objects whose "volume" (as described elsewhere in this section) is smaller than the given value. Spacetime objects whose volume is less than this will not participate in the matching and merging process, and no attribute information will be written to the ASCII output files. The default value is 10,000. If this seems rather large, consider the following example: Suppose the user is running MTD on a :math:`600 \times 400` grid, using 24 time steps. Then the volume of the whole data field is 600 :math:`\times` 400 :math:`\times` 24 = 5,760,000 cells. An object of volume 10,000 represents only 10,000/5,760,000 = 1/576 of the total data field. Setting **min\_volume** too small will typically produce a very large number of small objects, slowing down the MTD run and increasing the size of the output files.The configuration options listed above are common to many MODE and are described in :numref:`MODE-configuration-file`. +The **min_volume** entry tells MTD to throw away objects whose "volume" (as described elsewhere in this section) is smaller than the given value. Spacetime objects whose volume is less than this will not participate in the matching and merging process, and no attribute information will be written to the ASCII output files. The default value is 2,000. If this seems rather large, consider the following example: Suppose the user is running MTD on a :math:`600 \times 400` grid, using 24 time steps. Then the volume of the whole data field is 600 :math:`\times` 400 :math:`\times` 24 = 5,760,000 cells. An object of volume 2,000 represents only 2,000/5,760,000 = 1/2,880 of the total data field. Setting **min\_volume** too small will typically produce a very large number of small objects, slowing down the MTD run and increasing the size of the output files. The configuration options listed above are common to MODE and are described in :numref:`MODE-configuration-file`. ______________________ @@ -582,7 +582,7 @@ The contents of the OBJECT_ID and OBJECT_CAT columns identify the objects using - Integer * - 36 - CDIST_TRAVELLED - - Total great circle distance travelled by the 2D spatial centroid over the lifetime of the 3D object (in kilometers) + - Total great circle distance traveled by the 2D spatial centroid over the lifetime of the 3D object (in kilometers) - Double * - 37-41 - INTENSITY_10,_25,_50,_75,_90 @@ -662,7 +662,7 @@ MTD writes a NetCDF file containing various types of information as specified in • **Latitude** and **longitude** of all the points in the 2D grid. Useful for geolocating points or regions given by grid coordinates. -• **Raw data** from the input data files. This can be useful if the input data were grib format, since NetCDF is often easier to read. +• **Raw data** from the input data files. This can be useful if the input data were GRIB format, since NetCDF is often easier to read. • **Object ID** numbers, giving for each grid point the number of the simple object (if any) that covers that point. These numbers are one-based. A value of zero means that this point is not part of any object. diff --git a/docs/Users_Guide/mode.rst b/docs/Users_Guide/mode.rst index a15b55b628..24ea144e42 100644 --- a/docs/Users_Guide/mode.rst +++ b/docs/Users_Guide/mode.rst @@ -55,7 +55,7 @@ An example of the steps involved in resolving objects is shown in :numref:`mode- .. figure:: figure/mode-object_id.png - Example of an application of the MODE object identification process to a model precipitation field. + Example of an application of the MODE object identification process to a model precipitation field. .. _mode-attributes: @@ -89,7 +89,7 @@ Once object attributes :math:`\alpha_1,\alpha_2,\ldots,\alpha_n` are estimated, The next step is to define confidence maps :math:`C_i` for each attribute. These maps (again with values ranging from zero to one) reflect how confident we are in the calculated value of an attribute. The confidence maps generally are functions of the entire attribute vector :math:`\alpha = (\alpha_1, \alpha_2, \ldots, \alpha_n)`, in contrast to the interest maps, where each :math:`I_i` is a function only of :math:`\alpha_i`. To see why this is necessary, imagine an electronic anemometer that outputs a stream of numerical values of wind speed and direction. It is typically the case for such devices that when the wind speed becomes small enough, the wind direction is poorly resolved. The wind must be at least strong enough to overcome friction and turn the anemometer. Thus, in this case, our confidence in one attribute (wind direction) is dependent on the value of another attribute (wind speed). In MODE, all of the confidence maps except the map for axis angle are set to a constant value of 1. The axis angle confidence map is a function of aspect ratio, with values near one having low confidence, and values far from one having high confidence. -Next, scalar weights :math:`\boldsymbol{w}_i` are assigned to each attribute, representing an empirical judgment regarding the relative importance of the various attributes. As an example, the initial development of MODE, centroid distance was weighted more heavily than other attributes, because the location of storm systems close to each other in space seemed to be a strong indication (stronger than that given by any other attribute) that they were related. +Next, scalar weights :math:`\boldsymbol{w}_i` are assigned to each attribute, representing an empirical judgment regarding the relative importance of the various attributes. As an example, in the initial development of MODE, centroid distance was weighted more heavily than other attributes, because the location of storm systems close to each other in space seemed to be a strong indication (stronger than that given by any other attribute) that they were related. Finally, all these ingredients are collected into a single number called the total interest, :math:`\boldsymbol{T}`, given by: @@ -111,13 +111,13 @@ Multi-Variate MODE Traditionally, MODE defines objects by smoothing and thresholding data from a single input field. MET version 10.1.0 extends MODE by adding the option to define objects using multiple input fields. -As described in :numref:`MODE-configuration-file`, the **field** entry in the forecast and observation dictionaries define the input data to be processed. If **field** is defined as a dictionary, the traditional method for running MODE is invoked, where objects are defined using a single input field. If **field** is defined as an array of dictionaries, each specifying a different input field, then the multi-variate MODE logic is invoked and requires the **multivar_logic** configuration entry to be set. Traditional MODE is run once for each input field to define objects for that field. Note that the object definition criteria can be defined separately for each field array entry. The objects from each input field are combined into forecast and observation data *super* objects +As described in :numref:`MODE-configuration-file`, the **field** entry in the forecast and observation dictionaries define the input data to be processed. If **field** is defined as a dictionary, the traditional method for running MODE is invoked, where objects are defined using a single input field. If **field** is defined as an array of dictionaries, each specifying a different input field, then the multi-variate MODE logic is invoked and requires the **multivar_logic** configuration entry to be set. Traditional MODE is run once for each input field to define objects for that field. Note that the object definition criteria can be defined separately for each field array entry. The objects from each input field are combined into forecast and observation data *super* objects. The **multivar_logic** configuration entry, described in :numref:`MODE-configuration-file`, defines the boolean logic for combining objects from multiple fields into *super* objects. It can be defined once to apply to both the forecast and observation dictionaries if the field array lengths are the same, or defined separately within each dictionary. If defined separately within each dictionary, the field array lengths do not need to be the same for the forecast and observations. Note that the multi-variate MODE forecast and observation input fields and combination logic do not need to match. -The **multivar_intensity_compare_fcst** and **multivar_intensity_compare_obs** configuration entries, described in :numref:`MODE-configuration-file`, define the field array indexes for which to optionally compare intensities for individual input fields when the input is masked to non-missing only inside the *super* objects and are required to be the same length. For example, if **multivar_intensity_compare_fcst = [ 1, 2 ];** and **multivar_intensity_compare_obs = [ 2, 3 ];**, then index 1 (2) of the forecast field array will be compared with index 2 (3) of the observation field array. If an intensity comparision is requested, the corresponding pair of fields (fcst and obs) are masked to non-missing inside the fcst and obs super objects, and traditional mode is run on that pair of masked inputs producing uniquely named outputs. If **multivar_intensity_compare_fcst** and **multivar_intensity_compare_obs** are empty, the forecast and observation *super* objects are written to NetCDF, text, and postscript output files in the standard mode output format, but with no intensity information. +The **multivar_intensity_compare_fcst** and **multivar_intensity_compare_obs** configuration entries, described in :numref:`MODE-configuration-file`, define the field array indexes for which to optionally compare intensities for individual input fields when the input is masked to non-missing only inside the *super* objects and are required to be the same length. For example, if **multivar_intensity_compare_fcst = [ 1, 2 ];** and **multivar_intensity_compare_obs = [ 2, 3 ];**, then index 1 (2) of the forecast field array will be compared with index 2 (3) of the observation field array. If an intensity comparison is requested, the corresponding pair of fields (fcst and obs) are masked to non-missing inside the fcst and obs super objects, and traditional mode is run on that pair of masked inputs producing uniquely named outputs. If **multivar_intensity_compare_fcst** and **multivar_intensity_compare_obs** are empty, the forecast and observation *super* objects are written to NetCDF, text, and PostScript output files in the standard mode output format, but with no intensity information. -When regridding to the FCST or OBS field (e.g. to_grid = FCST), the first field of the field array is used from the forecast and observation field dictionaries, respectively. All regridding is then done to that grid. Other regrid options described in :ref:`regrid` can also be used as normal. +When regridding to the FCST or OBS field (e.g., to_grid = FCST), the first field of the field array is used from the forecast and observation field dictionaries, respectively. All regridding is then done to that grid. Other regrid options described in :ref:`regrid` can also be used as normal. "file_type" can be set independently for each input in multivariate mode. If not set for an input, MET uses file names and file content to determine the type. @@ -128,16 +128,16 @@ In multivariate mode, with quilt=**TRUE**, for all inputs the number of forecast When setting a threshold to a percentile, some choices require both an observation input and a forecast input. When this is the case, it's assumed the indices match, so for example if forecast input 1 has such a percentile setting, then observation input 1 will be used to compute the percentile. Percentiles in which this will happen are: * SFP in an observation input. - * The matching forecast input will be used to determine the threshold. e.g. ">SFP33.3" in the 2nd observation input means greater than 33.3-rd percentile of the 2nd forecast input will be used as the threshold for that observation input. + * The matching forecast input will be used to determine the threshold. e.g., ">SFP33.3" in the 2nd observation input means greater than 33.3-rd percentile of the 2nd forecast input will be used as the threshold for that observation input. * SOP in a forecast input. - * The matching observation input will be used to determine the threshold. e.g. ">SOP33.3" in the 2nd forecast input means greater than 33.3-rd percentile of the 2nd observation input will be used as the threshold for that forecast input. + * The matching observation input will be used to determine the threshold. e.g., ">SOP33.3" in the 2nd forecast input means greater than 33.3-rd percentile of the 2nd observation input will be used as the threshold for that forecast input. * "==FBIAS" in an observation input. - * e.g. "==FBIAS1" in an observation input to automatically de-bias the data, using a simple threshold in the matching forecast input. For example, when observation input 3 has "==FBIAS1", and forecast input 3 has ">5.0", MET applies the >5.0 threshold to the forecast and then chooses an observation threshold which results in a frequency bias of 1. The frequency bias can be any float value > 0.0. + * e.g., "==FBIAS1" in an observation input to automatically de-bias the data, using a simple threshold in the matching forecast input. For example, when observation input 3 has "==FBIAS1", and forecast input 3 has ">5.0", MET applies the >5.0 threshold to the forecast and then chooses an observation threshold which results in a frequency bias of 1. The frequency bias can be any float value > 0.0. * "==FBIAS" in a forecast input. - * e.g. "==FBIAS1" in a forecast input to automatically de-bias the data, using a simple threshold in the matching observation input. For example, when forecast input 2 has "==FBIAS1", and observation input 2 has ">5.0", MET applies the >5.0 threshold to the observation and then chooses a forecast threshold which results in a frequency bias of 1. The frequency bias can be any float value > 0.0. + * e.g., "==FBIAS1" in a forecast input to automatically de-bias the data, using a simple threshold in the matching observation input. For example, when forecast input 2 has "==FBIAS1", and observation input 2 has ">5.0", MET applies the >5.0 threshold to the observation and then chooses a forecast threshold which results in a frequency bias of 1. The frequency bias can be any float value > 0.0. Practical Information @@ -221,7 +221,7 @@ The default configuration file for the MODE tool, **MODEConfig_default**, can be A second default configuration file for the multivar MODE option, **MODEMultivarConfig_default**, is also found in the installed *share/met/config* directory. We encourage users to make a copy of this default configuration file when setting up a multivar configuration prior to modifying content. The two default config files **MODEConfig_default** and **MODEMultivarConfig_default** are similar, with **MODEMultivarConfig_default** having example multivar specific content. -Note that environment variables may be used when editing configuration files, as described in the :numref:`config_env_vars`. +Note that environment variables may be used when editing configuration files, as described in :numref:`config_env_vars`. _____________________ @@ -261,7 +261,7 @@ _____________________ .. code-block:: none - multivar_logic = "#1 && #2 && #3"; + multivar_logic = "#1 && #2 && #3"; The **multivar_logic** entry appears only in the **MODEMultivarConfig_default** file. This option applies to running multi-variate MODE by setting **field** to an array of dictionaries to define multiple input fields. Objects are defined separately for each input field based on the configuration settings specified for each field array entry. The **multivar_logic** entry is a string which defines how objects for each field are combined into a final *super* object. The objects for each field are referred to as '#N' where N is the N-th field array entry. The '&&' and '||' strings define intersection and union logic, respectively. For example, "#1 && #2" is the intersection of the objects from the first and second fields. "(#1 && #2) || #3" is the union of that intersection with the objects from the third field. @@ -271,8 +271,8 @@ _____________________ .. code-block:: none - multivar_intensity_compare_fcst = [1,2]; - multivar_intensity_compare_obs = [2,3]; + multivar_intensity_compare_fcst = [1,2]; + multivar_intensity_compare_obs = [2,3]; The **multivar_intensity_compare_fcst** and **multivar_intensity_compare_obs** entries appear only in the **MODEMultivarConfig_default** file. These entries define an index in the field arrays to be compared for forecast and observation intensities and must be the same length. For example, in the above example, forecast field 1 will be compared to observation field 2 for computing intensity attribute statistics. If the **multivar_intensity_compare_fcst** and **multivar_intensity_compare_obs** are empty, traditional mode output is created for the super objects, but with no intensity information. @@ -280,17 +280,17 @@ _____________________ .. code-block:: none - multivar_name = "Super"; + multivar_name = "Super"; -The **multivar_name** entry appears only in the **MODEMultivarConfig_default** file. This option is used only when the multivar option is enabled, and only when **multivar_intensity_compare_fcst** and **multivar_intensity_compare_obs** are empty. It can be thought of as an identifier for the multivariate super object. It shows up in output files names and content. It can be set separately for forecasts and observations or as a common value for both. +The **multivar_name** entry appears only in the **MODEMultivarConfig_default** file. This option is used only when the multivar option is enabled, and only when **multivar_intensity_compare_fcst** and **multivar_intensity_compare_obs** are empty. It can be thought of as an identifier for the multivariate super object. It shows up in output file names and content. It can be set separately for forecasts and observations or as a common value for both. _____________________ .. code-block:: none - multivar_level = "LO"; + multivar_level = "LO"; -The **multivar_level** entry appears only in the **MODEMultivarConfig_default** file. This option is used only when the multivar option is enabled, and only when **multivar_intensity_compare_fcst** and **multivar_intensity_compare_obs** are empty. It is the identifier for the multivariate super object as regards level. It shows up in output files names and content. If not set the default value is "NA". It can be set separately for forecasts and observations, or as a common value for both. +The **multivar_level** entry appears only in the **MODEMultivarConfig_default** file. This option is used only when the multivar option is enabled, and only when **multivar_intensity_compare_fcst** and **multivar_intensity_compare_obs** are empty. It is the identifier for the multivariate super object as regards level. It shows up in output file names and content. If not set the default value is "NA". It can be set separately for forecasts and observations, or as a common value for both. _____________________ @@ -313,15 +313,15 @@ _____________________ } obs = fcst; -The **field** entries in the forecast and observation dictionaries specify the model and observation variables and level to be compared. See a more complete description of them in :numref:`config_options`. In the above example, the forecast settings are copied into the observation dictionary using **obs = fcst;.** +The **field** entries in the forecast and observation dictionaries specify the model and observation variables and level to be compared. See a more complete description of them in :numref:`config_options`. In the above example, the forecast settings are copied into the observation dictionary using **obs = fcst;**. When **field** is set to an array of dictionaries rather than a single one, the multi-variate MODE logic is invoked. Please see :numref:`MODE-multivar` for a description of that logic. The **censor_thresh** and **censor_val** entries are used to censor the raw data as described in :numref:`config_options`. Their functionality replaces the **raw_thresh** entry, which is deprecated in met-6.1. Prior to defining objects, it is recommended that the raw fields should be made to look similar to each other. For example, if the model only predicts values for a variable above some threshold, the observations should be thresholded at that same level. The censor thresholds can be specified using symbols. By default, no censor thresholding is applied. -The **conv_radius** entry defines the radius of the circular convolution applied to smooth the raw fields. The radii are specified in terms of grid units. The default convolution radii are defined in terms of the previously defined **grid_res** entry. Multiple convolution radii may be specified as an array (e.g. **conv_radius = [ 5, 10, 15 ];**). +The **conv_radius** entry defines the radius of the circular convolution applied to smooth the raw fields. The radii are specified in terms of grid units. The default convolution radii are defined in terms of the previously defined **grid_res** entry. Multiple convolution radii may be specified as an array (e.g., **conv_radius = [ 5, 10, 15 ];**). -The **conv_thresh** entry specifies the threshold values to be applied to the convolved field to define objects. By default, objects are defined using a convolution threshold of 5.0. Multiple convolution thresholds may be specified as an array (e.g. **conv_thresh = [ >=5.0, >=10.0, >=15.0 ];)**. +The **conv_thresh** entry specifies the threshold values to be applied to the convolved field to define objects. By default, objects are defined using a convolution threshold of 5.0. Multiple convolution thresholds may be specified as an array (e.g., **conv_thresh = [ >=5.0, >=10.0, >=15.0 ];**). Multiple convolution radii and thresholds are processed using the logic defined by the **quilt** entry. The logic specific to multivariate mode is described in the multivariate mode section above. @@ -335,7 +335,7 @@ The **filter_attr_thresh** entry is an array of thresholds for these object attr Note that the **area_thresh** and **inten_perc_thresh** entries from earlier versions of MODE are replaced by these options and are now deprecated. -The **merge_thresh** entry is used to define larger objects for use in merging the original objects. It defines the threshold value used in the double thresholding merging technique. Note that in order to use this merging technique, it must be requested for both the forecast and observation fields. These thresholds should be chosen to define larger objects that fully contain the originally defined objects. For example, for objects defined as >=5.0, a merge threshold of >=2.5 will define larger objects that fully contain the original objects. Any two original objects contained within the same larger object will be merged. By default, the merge thresholds are set to be greater than or equal to 1.25. Multiple merge thresholds may be specified as an array (e.g. **merge_thresh = [ >=1.0, >=2.0, >=3.0 ];**). The number of **merge_thresh** entries must match the number of **conv_thresh** entries. +The **merge_thresh** entry is used to define larger objects for use in merging the original objects. It defines the threshold value used in the double thresholding merging technique. Note that in order to use this merging technique, it must be requested for both the forecast and observation fields. These thresholds should be chosen to define larger objects that fully contain the originally defined objects. For example, for objects defined as >=5.0, a merge threshold of >=2.5 will define larger objects that fully contain the original objects. Any two original objects contained within the same larger object will be merged. By default, the merge thresholds are set to be greater than or equal to 1.25. Multiple merge thresholds may be specified as an array (e.g., **merge_thresh = [ >=1.0, >=2.0, >=3.0 ];**). The number of **merge_thresh** entries must match the number of **conv_thresh** entries. The **merge_flag** entry controls what type of merging techniques will be applied to the objects defined in each field. @@ -455,7 +455,7 @@ _____________________ corner = 0.8; ratio_if = ( ( 0.0, 0.0 ) ( corner, 1.0 ) - ( 1.0, 1.0 ) ); + ( 1.0, 1.0 ) ); area_ratio = ratio_if; int_area_ratio = ( ... ); curvature_ratio = ratio_if; @@ -501,7 +501,7 @@ _____________________ color_table = "MET_BASE/colortables/mode_obj.ctable"; } -Specifying dictionaries to define the **color_table, plot_min**, and **plot_max** entries are described in :numref:`config_options`. +Specifying dictionaries to define the **color_table, plot_min**, and **plot_max** entries is described in :numref:`config_options`. The MODE tool generates a color bar to represent the contents of the colortable that was used to plot a field of data. The number of entries in the color bar matches the number of entries in the color table. The values defined for each color in the color table are also plotted next to the color bar. @@ -512,7 +512,7 @@ _____________________ plot_valid_flag = FALSE; -When applied, the **plot_valid_flag entry** indicates that only the region containing valid data after masking is applied should be plotted. +When applied, the **plot_valid_flag** entry indicates that only the region containing valid data after masking is applied should be plotted. • **FALSE** indicates the entire domain should be plotted. @@ -564,7 +564,7 @@ _____________________ shift_right = 0; -When MODE is run on global grids, this parameter specifies how many grid squares to shift the grid to the right. MODE does not currently connect objects from one side of a global grid to the other, potentially causing objects straddling the "cut" longitude to be separated into two objects. Shifting the grid by integer number of grid units enables the user to control where that longitude cut line occurs. +When MODE is run on global grids, this parameter specifies how many grid squares to shift the grid to the right. MODE does not currently connect objects from one side of a global grid to the other, potentially causing objects straddling the "cut" longitude to be separated into two objects. Shifting the grid by an integer number of grid units enables the user to control where that longitude cut line occurs. .. _MODE-output: @@ -784,15 +784,15 @@ This first file uses the following naming convention: *mode\_PREFIX\_FCST\_VAR\_LVL\_vs\_OBS\_VAR\_LVL\_HHMMSSL\_YYYYMMDD\_HHMMSSV\_HHMMSSA\_cts.txt* -where *PREFIX* indicates the user-defined output prefix, *FCST\_VAR\_LVL* is the forecast variable and vertical level being used, *OBS\_VAR\_LVL* is the observation variable and vertical level being used, *HHMMSSL* indicates the forecast lead time, *YYYYMMDD\_HHMMSSV* indicates the forecast valid time, and *HHMMSSA* indicates the accumulation period. The {\tt cts} string stands for contingency table statistics. The generation of this file can be disabled using the *ct\_stats\_flag* option in the configuration file. This CTS output file differs somewhat from the CTS output of the Point-Stat and Grid-Stat tools. The columns of this output file are summarized in :numref:`CTS_output`. +where *PREFIX* indicates the user-defined output prefix, *FCST\_VAR\_LVL* is the forecast variable and vertical level being used, *OBS\_VAR\_LVL* is the observation variable and vertical level being used, *HHMMSSL* indicates the forecast lead time, *YYYYMMDD\_HHMMSSV* indicates the forecast valid time, and *HHMMSSA* indicates the accumulation period. The *cts* string stands for contingency table statistics. The generation of this file can be disabled using the *ct\_stats\_flag* option in the configuration file. This CTS output file differs somewhat from the CTS output of the Point-Stat and Grid-Stat tools. The columns of this output file are summarized in :numref:`CTS_output`. -The second ASCII file the MODE tool generates contains all of the attributes for simple objects, the merged cluster objects, and pairs of objects. Each line in this file contains the same number of columns, though those columns not applicable to a given line contain fill data. The first row of every MODE object attribute file is a header containing the column names. The number of lines in this file depends on the number of objects defined. This file contains lines of 6 types that are indicated by the contents of the **OBJECT_ID** column. The **OBJECT_ID** can take the following 6 forms: **FNN, ONN, FNNN_ONNN, CFNNN, CONNN, CFNNN_CONNN**. In each case, **NNN** is a three-digit number indicating the object index. While all lines have the first 18 header columns in common, these 6 forms for **OBJECT_ID** can be divided into two types - one for single objects and one for pairs of objects. The single object lines **(FNN, ONN, CFNNN**, and **CONNN)** contain valid data in columns 19-39 and fill data in columns 40-51. The object pair lines **(FNNN_ONNN** and **CFNNN_CONNN)** contain valid data in columns 40-51 and fill data in columns 19-39. +The second ASCII file the MODE tool generates contains all of the attributes for simple objects, the merged cluster objects, and pairs of objects. Each line in this file contains the same number of columns, though those columns not applicable to a given line contain fill data. The first row of every MODE object attribute file is a header containing the column names. The number of lines in this file depends on the number of objects defined. This file contains lines of 6 types that are indicated by the contents of the **OBJECT_ID** column. The **OBJECT_ID** can take the following 6 forms: **FNNN, ONNN, FNNN_ONNN, CFNNN, CONNN, CFNNN_CONNN**. In each case, **NNN** is a three-digit number indicating the object index. While all lines have the first 18 header columns in common, these 6 forms for **OBJECT_ID** can be divided into two types - one for single objects and one for pairs of objects. The single object lines **(FNNN, ONNN, CFNNN**, and **CONNN)** contain valid data in columns 19-39 and fill data in columns 40-51. The object pair lines **(FNNN_ONNN** and **CFNNN_CONNN)** contain valid data in columns 40-51 and fill data in columns 19-39. These object identifiers are described in :numref:`MODE_object_attribute`. .. role:: raw-html(raw) - :format: html + :format: html .. _MODE_object_attribute: @@ -964,7 +964,7 @@ The contents of the columns in this ASCII file are summarized in :numref:`MODE_o - Double * - 37 - COMPLEXITY - - Ratio of the difference between the area of an object and the area of its convex hull divided by the area of the complex hull (unitless) + - Ratio of the difference between the area of an object and the area of its convex hull divided by the area of the convex hull (unitless) - Double * - 38-42 - INTENSITY_10, _25, _50, _75, _90 @@ -980,7 +980,7 @@ The contents of the columns in this ASCII file are summarized in :numref:`MODE_o - Double * - 45 - CENTROID_DIST - - Distance between two objects centroids (in grid units) + - Distance between two objects' centroids (in grid units) - Double * - 46 - BOUNDARY_DIST @@ -1051,9 +1051,9 @@ The dimensions and variables included in the mode NetCDF files are described in * - NetCDF Dimension - Description * - lat - - Dimension of the latitude (i.e. Number of grid points in the North-South direction) + - Dimension of the latitude (i.e., Number of grid points in the North-South direction) * - lon - - Dimension of the longitude (i.e. Number of grid points in the East-West direction) + - Dimension of the longitude (i.e., Number of grid points in the East-West direction) * - fcst_thresh_length - Number of thresholds applied to the forecast * - obs_thresh_length @@ -1083,7 +1083,7 @@ The dimensions and variables included in the mode NetCDF files are described in .. _Variables_contained_in_MODE_NetCDF_output: .. role:: raw-html(raw) - :format: html + :format: html .. list-table:: Variables contained in MODE NetCDF output. :widths: auto @@ -1306,7 +1306,7 @@ The dimensions and variables included in the mode NetCDF files are described in - Observation Cluster Convex Hull Point Y-Coordinate - Integer -**Postscript File** +**PostScript File** Lastly, the MODE tool creates a PostScript plot summarizing the features-based approach used in the verification. The PostScript plot is generated using internal libraries and does not depend on an external plotting package. The generation of this PostScript output can be disabled using the **ps_plot_flag** configuration file option. diff --git a/docs/Users_Guide/overview.rst b/docs/Users_Guide/overview.rst index c9956e2bc0..9c02f4ff82 100644 --- a/docs/Users_Guide/overview.rst +++ b/docs/Users_Guide/overview.rst @@ -9,7 +9,7 @@ Purpose and Organization of the User's Guide The goal of this User's Guide is to provide basic information for users of the Model Evaluation Tools (MET) to enable them to apply MET to their datasets and evaluation studies. MET was originally designed for application to the post-processed output of the `Weather Research and Forecasting (WRF) `_ model. However, MET may also be used for the evaluation of forecasts from other models or applications, including the `Unified Forecast System (UFS) `_, and the `System for Integrated Modeling of the Atmosphere (SIMA) `_ if certain file format definitions (described in this document) are followed. -The MET User's Guide is organized as follows. :numref:`overview` provides an overview of MET and its components. :numref:`installation` contains basic information about how to get started with MET - including system requirements, required software (and how to obtain it), how to download MET, and information about compilers, libraries, and how to build the code. :numref:`data_io` - :numref:`masking` focuses on the data needed to run MET, including formats for forecasts, observations, and output. These sections also document the reformatting and masking tools available in MET. :numref:`point-stat` - :numref:`gsi_tools` focuses on the main statistics modules contained in MET, including the Point-Stat, Grid-Stat, Ensemble-Stat, Wavelet-Stat and GSI Diagnostic Tools. These sections include an introduction to the statistical verification methodologies utilized by the tools, followed by a section containing practical information, such as how to set up configuration files and the format of the output. :numref:`stat-analysis` and :numref:`series-analysis` focus on the analysis modules, Stat-Analysis and Series-Analysis, which aggregate the output statistics from the other tools across multiple cases. :numref:`mode` - :numref:`mode-td` describes a suite of object-based tools, including MODE, MODE-Analysis, and MODE-TD. :numref:`met-tc_overview` - :numref:`rmw-analysis` describes tools focused on tropical cyclones, including MET-TC Overview, TC-DLand, TC-Diag, TC-Pairs, TC-Stat, TC-Gen, TC-RMW and RMW-Analysis. Finally, :numref:`plotting` includes plotting tools included in the MET release for checking and visualizing data, as well as some additional tools and information for plotting MET results. The appendices provide further useful information, including answers to some typical questions (:numref:`Appendix A, Section %s `) and links and information about map projections, grids, and polylines (:numref:`Appendix B, Section %s `). :numref:`Appendix C, Section %s ` and :numref:`Appendix D, Section %s ` provide more information about the verification measures and confidence intervals that are provided by MET. Sample code that can be used to perform analyses on the output of MET and create particular types of plots of verification results is posted on the `MET website `_). Note that the MET development group also accepts contributed analysis and plotting scripts which may be posted on the MET website for use by the community. It should be noted there are References (:numref:`refs`) in this User's Guide as well. +The MET User's Guide is organized as follows. :numref:`overview` provides an overview of MET and its components. :numref:`installation` contains basic information about how to get started with MET - including system requirements, required software (and how to obtain it), how to download MET, and information about compilers, libraries, and how to build the code. :numref:`data_io` - :numref:`masking` focuses on the data needed to run MET, including formats for forecasts, observations, and output. These sections also document the reformatting and masking tools available in MET. :numref:`point-stat` - :numref:`gsi_tools` focuses on the main statistics modules contained in MET, including the Point-Stat, Grid-Stat, Ensemble-Stat, Wavelet-Stat and GSI Diagnostic Tools. These sections include an introduction to the statistical verification methodologies utilized by the tools, followed by a section containing practical information, such as how to set up configuration files and the format of the output. :numref:`stat-analysis` and :numref:`series-analysis` focus on the analysis modules, Stat-Analysis and Series-Analysis, which aggregate the output statistics from the other tools across multiple cases. :numref:`mode` - :numref:`mode-td` describes a suite of object-based tools, including MODE, MODE-Analysis, and MODE-TD. :numref:`met-tc_overview` - :numref:`rmw-analysis` describes tools focused on tropical cyclones, including MET-TC Overview, TC-DLand, TC-Diag, TC-Pairs, TC-Stat, TC-Gen, TC-RMW and RMW-Analysis. Finally, :numref:`plotting` includes plotting tools included in the MET release for checking and visualizing data, as well as some additional tools and information for plotting MET results. The appendices provide further useful information, including answers to some typical questions (:numref:`Appendix A, Section %s `) and links and information about map projections, grids, and polylines (:numref:`Appendix B, Section %s `). :numref:`Appendix C, Section %s ` and :numref:`Appendix D, Section %s ` provide more information about the verification measures and confidence intervals that are provided by MET. Sample code that can be used to perform analyses on the output of MET and create particular types of plots of verification results is posted on the `MET website `_. Note that the MET development group also accepts contributed analysis and plotting scripts which may be posted on the MET website for use by the community. It should be noted there are References (:numref:`refs`) in this User's Guide as well. The remainder of this section includes information about the context for MET development, as well as information on the design principles used in developing MET. In addition, this section includes an overview of the MET package and its specific modules. @@ -25,11 +25,11 @@ The MET package is available to DTC staff, visitors, and collaborators, as well MET Goals and Design Philosophy =============================== -The primary goal of MET development is to provide a state-of-the-art verification package to the NWP community. By "state-of-the-art" we mean that MET will incorporate newly developed and advanced verification methodologies, including new methods for diagnostic and spatial verification and new techniques provided by the verification and modeling communities. MET also utilizes and replicates the capabilities of existing systems for verification of NWP forecasts. For example, the MET package replicates existing National Center for Environmental Prediction (NCEP) operational verification capabilities (e.g., I/O, methods, statistics, data types). MET development will take into account the needs of the NWP community - including operational centers and the research and development community. Some of the MET capabilities include traditional verification approaches for standard surface and upper air variables (e.g., Equitable Threat Score, Mean Squared Error), confidence intervals for verification measures, and spatial forecast verification methods. In the future, MET will include additional state-of-the-art and new methodologies. +The primary goal of MET development is to provide a state-of-the-art verification package to the NWP community. By "state-of-the-art" we mean that MET will incorporate newly developed and advanced verification methodologies, including new methods for diagnostic and spatial verification and new techniques provided by the verification and modeling communities. MET also utilizes and replicates the capabilities of existing systems for verification of NWP forecasts. For example, the MET package replicates existing National Centers for Environmental Prediction (NCEP) operational verification capabilities (e.g., I/O, methods, statistics, data types). MET development will take into account the needs of the NWP community - including operational centers and the research and development community. Some of the MET capabilities include traditional verification approaches for standard surface and upper air variables (e.g., Equitable Threat Score, Mean Squared Error), confidence intervals for verification measures, and spatial forecast verification methods. In the future, MET will include additional state-of-the-art and new methodologies. The MET package has been designed to be modular and adaptable. For example, individual modules can be applied without running the entire set of tools. New tools can easily be added to the MET package due to this modular design. In addition, the tools can readily be incorporated into a larger "system" that may include a database as well as more sophisticated input/output and user interfaces. Currently, the MET package is a set of tools that can easily be applied by any user on their own computer platform. A suite of Python scripts for low-level automation of verification workflows and plotting has been developed to assist users with setting up their MET-based verification. It is called METplus and may be obtained on the `METplus GitHub repository `_. -The MET code and documentation is maintained by the DTC in Boulder, Colorado. The MET package is freely available to the modeling, verification, and operational communities, including universities, governments, the private sector, and operational modeling and prediction centers. +The MET code and documentation are maintained by the DTC in Boulder, Colorado. The MET package is freely available to the modeling, verification, and operational communities, including universities, governments, the private sector, and operational modeling and prediction centers. MET Components ============== @@ -42,9 +42,9 @@ The reformatting stage of MET consists of several tools which perform a variety .. figure:: figure/overview-figure.png - Basic representation of current MET structure and modules. Gray areas represent input and output files. Dark green areas represent point preprocessing tools. Light green areas represent grid preprocessing tools. Pink areas represent regridding tools. Orange areas represent plotting tools. Blue areas represent statistical tools. Yellow areas represent aggregation and analysis tools. + Basic representation of current MET structure and modules. Gray areas represent input and output files. Dark green areas represent point preprocessing tools. Light green areas represent grid preprocessing tools. Pink areas represent regridding tools. Orange areas represent plotting tools. Blue areas represent statistical tools. Yellow areas represent aggregation and analysis tools. -Several optional plotting utilities are provided to assist users in checking their output from the data preprocessing step. Plot-Point-Obs creates a postscript plot showing the locations of point observations. This can be quite useful for assessing whether the latitude and longitude of observation stations was specified correctly. Plot-Data-Plane produces a similar plot for gridded data. Finally, WWMCA-Plot produces a plot of the raw WWMCA data file. +Several optional plotting utilities are provided to assist users in checking their output from the data preprocessing step. Plot-Point-Obs creates a postscript plot showing the locations of point observations. This can be quite useful for assessing whether the latitude and longitude of observation stations were specified correctly. Plot-Data-Plane produces a similar plot for gridded data. Finally, WWMCA-Plot produces a plot of the raw WWMCA data file. The main statistical analysis components of the current version of MET are: Point-Stat, Pair-Stat, Grid-Stat, Series-Analysis, Ensemble-Stat, MODE, MODE-TD (MTD), Grid-Diag, and Wavelet-Stat. @@ -56,11 +56,11 @@ Sometimes it may be useful to verify a forecast against gridded fields (e.g., St Users wishing to accumulate statistics over a time, height, or other series separately for each grid location should use the Series-Analysis tool. Series-Analysis can read any gridded matched pair data produced by the other MET tools and accumulate them, keeping each spatial location separate. Maps of these statistics can be useful for diagnosing spatial differences in forecast quality. -Ensemble-Stat compares ensemble member data to gridded analyses and/or point observations and computes measures of ensemble characteristics. The ensemble characteristics include ensemble mean and spread information, computation of rank and probability integral transform (PIT) histograms, the points for the receiver operator characteristic (ROC) and reliability diagrams, and ranked probabilities scores (RPS) and the continuous version (CRPS). When categorical thresholds are specified, Ensemble-Stat derives ensemble relative frequencies and verifies them as probability forecasts against the gridded analyses and/or point observations provided. Note that the ensemble post-processing provided in prior versions of this tool has moved to Gen-Ens-Prod. +Ensemble-Stat compares ensemble member data to gridded analyses and/or point observations and computes measures of ensemble characteristics. The ensemble characteristics include ensemble mean and spread information, computation of rank and probability integral transform (PIT) histograms, the points for the receiver operator characteristic (ROC) and reliability diagrams, and ranked probability scores (RPS) and the continuous version (CRPS). When categorical thresholds are specified, Ensemble-Stat derives ensemble relative frequencies and verifies them as probability forecasts against the gridded analyses and/or point observations provided. Note that the ensemble post-processing provided in prior versions of this tool has moved to Gen-Ens-Prod. The MODE (Method for Object-based Diagnostic Evaluation) tool also uses gridded fields as observational datasets. However, unlike the Grid-Stat tool, which applies traditional forecast verification techniques, MODE applies the object-based spatial verification technique described in :ref:`Davis et al. (2006a,b) ` and :ref:`Brown et al. (2007) `. This technique was developed in response to the "double penalty" problem in forecast verification. A forecast missed by even a small distance is effectively penalized twice by standard categorical verification scores: once for missing the event and a second time for producing a false alarm of the event elsewhere. As an alternative, MODE defines objects in both the forecast and observation fields. The objects in the forecast and observation fields are then matched and compared to one another. Applying this technique also provides diagnostic verification information that is difficult or even impossible to obtain using traditional verification measures. For example, the MODE tool can provide information about errors in location, size, and intensity. -The MODE-TD tool extends object-based analysis from two-dimensional forecasts and observations to include the time dimension. In addition to the two dimensional information provided by MODE, MODE-TD can be used to examine even more features including displacement in time, and duration and speed of moving areas of interest. +The MODE-TD tool extends object-based analysis from two-dimensional forecasts and observations to include the time dimension. In addition to the two-dimensional information provided by MODE, MODE-TD can be used to examine even more features including displacement in time, and duration and speed of moving areas of interest. The Grid-Diag tool produces multivariate probability density functions (PDFs) that may be used either for exploring the relationship between two fields, or for the computation of percentiles generated from the sample for use with percentile thresholding. The output from this tool requires post-processing by METplus or user-provided utilities. @@ -68,23 +68,23 @@ The Wavelet-Stat tool decomposes two-dimensional forecasts and observations acco Results from the statistical analysis stage are output in ASCII, NetCDF and Postscript formats. The Point-Stat, Grid-Stat, Wavelet-Stat, and Ensemble-Stat tools create STAT (statistics) files which are tabular ASCII files ending with a ".stat" suffix. The STAT output files consist of multiple line types, each containing a different set of related statistics. The columns preceding the LINE_TYPE column are common to all lines. However, the number and contents of the remaining columns vary by line type. -The Stat-Analysis and MODE-Analysis tools aggregate the output statistics from the previous steps across multiple cases. The Stat-Analysis tool reads the STAT output of Point-Stat, Grid-Stat, Ensemble-Stat, and Wavelet-Stat and can be used to filter the STAT data and produce aggregated continuous and categorical statistics. Stat-Analysis also reads matched pair data (i.e. MPR line type) via python embedding. The MODE-Analysis tool reads the ASCII output of the MODE tool and can be used to produce summary information about object location, size, and intensity (as well as other object characteristics) across one or more cases. +The Stat-Analysis and MODE-Analysis tools aggregate the output statistics from the previous steps across multiple cases. The Stat-Analysis tool reads the STAT output of Point-Stat, Grid-Stat, Ensemble-Stat, and Wavelet-Stat and can be used to filter the STAT data and produce aggregated continuous and categorical statistics. Stat-Analysis also reads matched pair data (i.e., MPR line type) via python embedding. The MODE-Analysis tool reads the ASCII output of the MODE tool and can be used to produce summary information about object location, size, and intensity (as well as other object characteristics) across one or more cases. -Tropical cyclone forecasts and observations are quite different than numerical model forecasts, and thus they have their own set of tools. These consist of TC-DLand, TC-Diag, TC-Pairs, TC-Stat, TC-Gen, TC-RMW, and RMW-Analysis. The TC-DLand module calculates the distance to land from all locations on a specified grid. This information can be used in later modules to eliminate tropical cyclones that are over land from being included in the statistics. TC-Diag converts gridded model output into cylindrical coordinates for each storm location, calls Python scripts to compute storm-relative diagnostics, and writes ASCII output to be read by TC-Pairs. TC-Pairs matches up tropical cyclone forecasts and observations and writes all output to a file. In TC-Stat, these forecast / observation pairs are analyzed according to user preference to produce statistics. TC-Gen evaluates the performance of Tropical Cyclone genesis forecast using contingency table counts and statistics. TC-RMW performs a coordinate transformation for gridded model or analysis fields centered on the current storm location. RMW-Analysis filters and aggregates the output of TC-RMW across multiple cases. +Tropical cyclone forecasts and observations are quite different than numerical model forecasts, and thus they have their own set of tools. These consist of TC-DLand, TC-Diag, TC-Pairs, TC-Stat, TC-Gen, TC-RMW, and RMW-Analysis. The TC-DLand module calculates the distance to land from all locations on a specified grid. This information can be used in later modules to eliminate tropical cyclones that are over land from being included in the statistics. TC-Diag converts gridded model output into cylindrical coordinates for each storm location, calls Python scripts to compute storm-relative diagnostics, and writes ASCII output to be read by TC-Pairs. TC-Pairs matches up tropical cyclone forecasts and observations and writes all output to a file. In TC-Stat, these forecast / observation pairs are analyzed according to user preference to produce statistics. TC-Gen evaluates the performance of Tropical Cyclone genesis forecasts using contingency table counts and statistics. TC-RMW performs a coordinate transformation for gridded model or analysis fields centered on the current storm location. RMW-Analysis filters and aggregates the output of TC-RMW across multiple cases. The following sections of this MET User's Guide contain usage statements for each tool, which may be viewed if you type the name of the tool. Alternatively, the user can also type the name of the tool followed by **-help** to obtain the usage statement. Each tool also has a **-version** command line option associated with it so that the user can determine what version of the tool they are using. Future Development Plans ======================== -MET is an evolving verification software package. New capabilities are provided in controlled, successive releases using sematic version numbering. Please refer to the issues listed in the `MET GitHub repository `_ to see our development priorities for upcoming releases. +MET is an evolving verification software package. New capabilities are provided in controlled, successive releases using semantic version numbering. Please refer to the issues listed in the `MET GitHub repository `_ to see our development priorities for upcoming releases. -Bugs and user-identified problems are documented as GitHub issues as they are found. Typically, bug fixes are provided in the next bugfix release for the current official realease as well as the next official release. Specific details about each bug can be found in the body and/or comments of the corresponding GitHub issue. +Bugs and user-identified problems are documented as GitHub issues as they are found. Typically, bug fixes are provided in the next bugfix release for the current official release as well as the next official release. Specific details about each bug can be found in the body and/or comments of the corresponding GitHub issue. User Support ============ -MET is the statistical component of the larger METplus system for which user support is provided through the `METplus GitHub Discussions Forum `_, as described in the `METplus User's Guide `_. Additional information about MET are provided on the `MET web page `_. +MET is the statistical component of the larger METplus system for which user support is provided through the `METplus GitHub Discussions Forum `_, as described in the `METplus User's Guide `_. Additional information about MET is provided on the `MET web page `_. .. note:: @@ -94,6 +94,6 @@ MET is the statistical component of the larger METplus system for which user sup Fortify and SonarQube ===================== -Requirements from various government agencies that use MET have resulted in our code being analyzed by both the Fortify and SonarQube static source code analysis tools. Fortify and SonarQube analyze source code to identify for security risks, memory leaks, uninitialized variables, and other such weaknesses and bad coding practices. They categorize issue as low priority, high priority, or critical, and report these issues back to the developers for them to address. The goal is to drive the counts of both high priority and critical issues down to zero. +Requirements from various government agencies that use MET have resulted in our code being analyzed by both the Fortify and SonarQube static source code analysis tools. Fortify and SonarQube analyze source code to identify security risks, memory leaks, uninitialized variables, and other such weaknesses and bad coding practices. They categorize issues as low priority, high priority, or critical, and report these issues back to the developers for them to address. The goal is to drive the counts of both high priority and critical issues down to zero. -The MET developers are pleased to report that Fortify reports zero critical issues in the MET code. Users of the MET tools who work in high security environments can rest assured about the possibility of security risks when using MET, since the quality of the code has now been vetted by unbiased third-party experts. The MET developers continue using Fortify routinely to ensure that the critical counts remain at zero and to further reduce the counts for lower priority issues. +The MET developers are pleased to report that Fortify reports zero critical issues in the MET code. Users of the MET tools who work in high security environments can rest assured that the risk of security issues when using MET is low, since the quality of the code has now been vetted by unbiased third-party experts. The MET developers continue using Fortify routinely to ensure that the critical counts remain at zero and to further reduce the counts for lower priority issues. diff --git a/docs/Users_Guide/pair-stat.rst b/docs/Users_Guide/pair-stat.rst index 76f351ac33..a400abb27f 100644 --- a/docs/Users_Guide/pair-stat.rst +++ b/docs/Users_Guide/pair-stat.rst @@ -47,7 +47,7 @@ Required Arguments for pair_stat 1. The **-pairs** argument defines one or more input files containing forecast/observation pairs. May be set as a list of file names (**file_1 ... file_n**) or as an ASCII file containing - a list file names (**file_list**), as described in :numref:`ascii_file_lists` (required). + a list of file names (**file_list**), as described in :numref:`ascii_file_lists` (required). This option can be used multiple times but all inputs must follow the same **-format type**, described below. @@ -101,7 +101,7 @@ The default configuration file for the Pair-Stat tool named **PairStatConfig_def in the installed *share/met/config* directory. Users are encouraged to make a copy prior to modifying its contents. The configuration file options are described in the subsections below. -Note that environment variables may be used when editing configuration files, as described in the +Note that environment variables may be used when editing configuration files, as described in :numref:`config_env_vars`. ________________________ @@ -150,7 +150,7 @@ entries in each array entry vary based on the input file format: desired values of the input :ref:`MET MPR Line Type`. Set **fcst.pairs.name** and **obs.pairs.name** to the desired values of the **FCST_VAR** and **OBS_VAR** columns, respectively. Set **fcst.pairs.level** and - and **obs.pairs.level** to one or more desired values of the **FCST_LEV** + **obs.pairs.level** to one or more desired values of the **FCST_LEV** and **OBS_LEV** columns, respectively. Only MPR lines whose variable names match those requested and whose level strings appear in the list of requested level strings will be used for that verification task. @@ -164,7 +164,7 @@ entries in each array entry vary based on the input file format: IODA files typically use NetCDF4 groups, and the **name** entry should specify both the group and variable names, formatted as ``name = "/GROUP_NAME/VARIABLE_NAME";`` - (e.g. ``name = "/hofx/air_temperature";``). + (e.g., ``name = "/hofx/air_temperature";``). _________________________ @@ -181,8 +181,8 @@ _________________________ The configuration options listed above are used to process the input paired data and define -thresholds for filtering data and for categorical verification. They can be speficied separately -for each verification task inside in each **fcst.pairs** or **obs.pairs** array entry. +thresholds for filtering data and for categorical verification. They can be specified separately +for each verification task inside each **fcst.pairs** or **obs.pairs** array entry. They are common to multiple MET tools and are described in :numref:`config_options`. _________________________ @@ -253,7 +253,7 @@ _________________________ The configuration options listed above filter the input paired data temporally. They can -be specified separately for each verificaiton task inside each **obs.pairs** array entry. +be specified separately for each verification task inside each **obs.pairs** array entry. They are also supported for the Stat-Analysis tool and are described in :numref:`stat_analysis-configuration-file`. _________________________ @@ -310,7 +310,7 @@ _________________________ The **mask** dictionary defines how the input paired data is aggregated spatially when computing statistics. For each verification task, output statistics will be computed for each masking region defined. This is -common to multiple MET tools and are described in :numref:`config_options`. +common to multiple MET tools and is described in :numref:`config_options`. .. note:: @@ -356,7 +356,7 @@ pair_stat Output ---------------- The Pair-Stat tool produces output in STAT and, optionally, ASCII format. The ASCII output -duplicates the STAT output but has the data organized by line type. The output files names +duplicates the STAT output but has the data organized by line type. The output file names begin with **./pair_stat** or the string specified with the **-out base** command line option. Users should set the **-out base** command line option appropriately to avoid overwriting output generated by previous runs of the Pair-Stat tool. diff --git a/docs/Users_Guide/plotting.rst b/docs/Users_Guide/plotting.rst index 239e23ae81..ccef4b23af 100644 --- a/docs/Users_Guide/plotting.rst +++ b/docs/Users_Guide/plotting.rst @@ -11,10 +11,10 @@ This section describes how to check your data files using plotting utilities. Po .. note:: - Earlier versions included the plot_mode_field utility to visualize the NetCDF output generated by the MODE tool. - This utility was deprecated and removed for MET version 13.0.0 to reduce MET's external library dependencies. - Users are encouraged to visualize the NetCDF output from MODE using the plot_data_plane utility and/or - using functionality provided in `METplotpy `_. + Earlier versions included the plot_mode_field utility to visualize the NetCDF output generated by the MODE tool. + This utility was deprecated and removed for MET version 13.0.0 to reduce MET's external library dependencies. + Users are encouraged to visualize the NetCDF output from MODE using the plot_data_plane utility and/or + using functionality provided in `METplotpy `_. plot_point_obs Usage -------------------- @@ -56,7 +56,7 @@ Optional Arguments for plot_point_obs 6. The **-title string** option specifies the plot title string. -7. The **-gc code** and **-obs_var name** options specify observation types to be plotted. These overrides the corresponding configuration file entries. +7. The **-gc code** and **-obs_var name** options specify observation types to be plotted. These override the corresponding configuration file entries. 8. The **-msg_typ name** option specifies the message type to be plotted. This overrides the corresponding configuration file entry. @@ -72,9 +72,9 @@ An example of the plot_point_obs calling sequence is shown below: plot_point_obs sample_pb.nc sample_data.ps -In this example, the Plot-Point-Obs tool will process the input sample_pb.nc file and write a postscript file containing a plot to a file named sample_pb.ps. +In this example, the Plot-Point-Obs tool will process the input sample_pb.nc file and write a PostScript file containing a plot to a file named sample_data.ps. -An equivalent command using python embedding for point observations is shown below. Note that the entire python command is enclosed in single quotes to prevent embedded whitespace for causing parsing errors: +An equivalent command using Python embedding for point observations is shown below. Note that the entire Python command is enclosed in single quotes to prevent embedded whitespace from causing parsing errors: .. code-block:: none @@ -87,7 +87,7 @@ plot_point_obs Configuration File The default configuration file for the Plot-Point-Obs tool named **PlotPointObsConfig_default** can be found in the installed *share/met/config* directory. The contents of the configuration file are described in the subsections below. -Note that environment variables may be used when editing configuration files, as described in the :numref:`config_env_vars`. +Note that environment variables may be used when editing configuration files, as described in :numref:`config_env_vars`. ________________________ @@ -125,7 +125,7 @@ The **grid_data** dictionary defines a gridded field of data to be plotted as a The **to_grid** entry in the **regrid** dictionary specifies if and how the requested gridded data should be regridded prior to plotting. Please see :numref:`config_options` for a description of the **regrid** dictionary options. -The **grid_plot_info** dictionary inside **grid_data** specifies the options for for plotting the gridded data. The options within **grid_plot_info** are described in :numref:`config_options`. +The **grid_plot_info** dictionary inside **grid_data** specifies the options for plotting the gridded data. The options within **grid_plot_info** are described in :numref:`config_options`. ______________________ @@ -147,7 +147,7 @@ ______________________ obs_var = []; obs_quality = []; -The options listed above define filtering criteria for the input point observation strings. If empty, no filtering logic is applied. If a comma-separated list of strings is provided, only those observations meeting all of the criteria are included. The **msg_typ** entry specifies the message type. The **sid_inc** and **sid_exc** entries explicitly specify station id's to be included or excluded. The **obs_var** entry specifies the observation variable names, and **obs_quality** specifies quality control strings. +The options listed above define filtering criteria for the input point observation strings. If empty, no filtering logic is applied. If a comma-separated list of strings is provided, only those observations meeting all of the criteria are included. The **msg_typ** entry specifies the message type. The **sid_inc** and **sid_exc** entries explicitly specify station IDs to be included or excluded. The **obs_var** entry specifies the observation variable names, and **obs_quality** specifies quality control strings. ______________________ @@ -193,7 +193,7 @@ ______________________ .. code-block:: none - dotsize(x) = 1.0; + dotsize(x) = 1.0; The **dotsize(x)** function defines the size of the circle to be plotted as a function of the observation value. The default setting shown above defines the dot size as a constant value. @@ -204,7 +204,7 @@ ______________________ line_color = []; line_width = 1; -The **line_color** and **line_width** entries define the color and thickness of the outline for each circle plotted. When **line_color** is left as an empty array, no outline is drawn. Otherwise, **line_color** should be specified using 3 intergers between 0 and 255 to define the red, green, and blue components of the color. +The **line_color** and **line_width** entries define the color and thickness of the outline for each circle plotted. When **line_color** is left as an empty array, no outline is drawn. Otherwise, **line_color** should be specified using 3 integers between 0 and 255 to define the red, green, and blue components of the color. ______________________ @@ -223,7 +223,7 @@ The circles are filled in based on the setting of the **fill_color** and **fill_ Users are encouraged to define as many **point_data** array entries as needed to filter and plot the input observations in the way they would like. Each point observation is plotted using the options specified in the first matching array entry. Note that the filtering, processing, and plotting options specified inside each **point_data** array entry take precedence over ones specified at the higher level of configuration file context. -For each observation, this tool stores the observation latitude, longitude, and value. However, unless the **dotsize(x)** function is not constant or the **fill_plot_info.flag** entry is set to true, the observation value is simply set to a flag value. For each **point_data** array entry, the tool stores and plots only the unique combination of observation latitude, longitude, and value. Therefore multiple obsevations at the same location will typically be plotted as a single circle. +For each observation, this tool stores the observation latitude, longitude, and value. However, unless the **dotsize(x)** function is not constant or the **fill_plot_info.flag** entry is set to true, the observation value is simply set to a flag value. For each **point_data** array entry, the tool stores and plots only the unique combination of observation latitude, longitude, and value. Therefore multiple observations at the same location will typically be plotted as a single circle. .. _plot_data_plane-usage: @@ -280,7 +280,7 @@ A second example of the plot_data_plane calling sequence is shown below: plot_data_plane test.grb2 test.ps 'name="DSWRF"; level="L0";' -v 4 -In the first example, the Plot-Data-Plane tool will process the input test.grb file and write a PostScript image to a file named test.ps showing temperature at 2 meters. The second example plots downward shortwave radiation flux at the surface. The second example is run at verbosity level 4 so that the user can inspect the output and make sure its plotting the intended record. +In the first example, the Plot-Data-Plane tool will process the input test.grb file and write a PostScript image to a file named test.ps showing temperature at 2 meters. The second example plots downward shortwave radiation flux at the surface. The second example is run at verbosity level 4 so that the user can inspect the output and make sure it's plotting the intended record. Examples of Plotting MET Output =============================== @@ -294,20 +294,20 @@ The plots in :numref:`plotting_Gilbert_skill_score` show time series of frequenc .. figure:: figure/plotting-Gilbert-skill-score.png - Time series of forecast area bias and Gilbert Skill Score for four model configurations (different lines) stratified by time-of-day. + Time series of forecast area bias and Gilbert Skill Score for four model configurations (different lines) stratified by time-of-day. MODE Tool Examples ------------------ When using the MODE tool, it is possible to think of matched objects as hits and unmatched objects as false alarms or misses depending on whether the unmatched object is from the forecast or observed field, respectively. Because the objects can have greatly differing sizes, it is useful to weight the statistics by the areas, which are given in the output as numbers of grid squares. When doing this, it is possible to have different matched observed object areas from matched forecast object areas so that the number of hits will be different depending on which is chosen to be a hit. When comparing multiple forecasts to the same observed field, it is perhaps wise to always use the observed field for the hits so that there is consistency for subsequent comparisons. Defining hits, misses and false alarms in this way allows one to compute many traditional verification scores without the problem of small-scale discrepancies; the matched objects are defined as being matched because they are "close" by the fuzzy logic criteria. Note that scores involving the number of correct negatives may be more difficult to interpret as it is not clear how to define a correct negative in this context. It is also important to evaluate the number and area attributes for these objects in order to provide a more complete picture of how the forecast is performing. -:numref:`plotting_verification` gives an example of two traditional verification scores (Bias and CSI) along with bar plots showing the total numbers of objects for the forecast and observed fields, as well as bar plots showing their total areas. These data are from the same set of 13-km WRF model runs analyzed in :numref:`plotting_verification`. The model runs were initialized at 0 UTC and cover the period 15 July to 15 August 2005. For the forecast evaluation, we compared 3-hour accumulated precipitation for lead times of 3-24 hours to Stage II radar-gauge precipitation. Note that for the 3-hr lead time, indicated as the 0300 UTC valid time in :numref:`plotting_Gilbert_skill_score`, the Bias is significantly larger than the other lead times. This is evidenced by the fact that there are both a larger number of forecast objects, and a larger area of forecast objects for this lead time, and only for this lead time. Dashed lines show about 2 bootstrap standard deviations from the estimate. +:numref:`plotting_verification` gives an example of two traditional verification scores (Bias and CSI) along with bar plots showing the total numbers of objects for the forecast and observed fields, as well as bar plots showing their total areas. These data are from the same set of 13-km WRF model runs analyzed in :numref:`plotting_Gilbert_skill_score`. The model runs were initialized at 0 UTC and cover the period 15 July to 15 August 2005. For the forecast evaluation, we compared 3-hour accumulated precipitation for lead times of 3-24 hours to Stage II radar-gauge precipitation. Note that for the 3-hr lead time, indicated as the 0300 UTC valid time in :numref:`plotting_verification`, the Bias is significantly larger than the other lead times. This is evidenced by the fact that there are both a larger number of forecast objects, and a larger area of forecast objects for this lead time, and only for this lead time. Dashed lines show about 2 bootstrap standard deviations from the estimate. .. _plotting_verification: .. figure:: figure/plotting-verification.png - Traditional verification scores applied to output of the MODE tool, computed by defining matched observed objects to be hits, unmatched observed objects to be misses, and unmatched forecast objects to be false alarms; weighted by object area. Bar plots show numbers (penultimate row) and areas (bottom row) of observed and forecast objects, respectively. + Traditional verification scores applied to output of the MODE tool, computed by defining matched observed objects to be hits, unmatched observed objects to be misses, and unmatched forecast objects to be false alarms; weighted by object area. Bar plots show numbers (penultimate row) and areas (bottom row) of observed and forecast objects, respectively. In addition to the traditional scores, MODE output allows more information to be gleaned about forecast performance. It is even useful when computing the traditional scores to understand how much the forecasts are displaced in terms of both distance and direction. :numref:`plotting_histogram`, for example, shows circle histograms for matched objects. The petals show the percentage of times the forecast object centroids are at a given angle from the observed object centroids. In :numref:`plotting_histogram` (top diagram) about 25% of the time the forecast object centroids are west of the observed object centroids, whereas in :numref:`plotting_histogram` (bottom diagram) there is less bias in terms of the forecast objects' centroid locations compared to those of the observed objects, as evidenced by the petals' relatively similar lengths, and their relatively even dispersion around the circle. The colors on the petals represent the proportion of centroid distances within each colored bin along each direction. For example, :numref:`plotting_histogram` (top row) shows that among the forecast object centroids that are located to the West of the observed object centroids, the greatest proportion of the separation distances (between the observed and forecast object centroids) is greater than 20 grid squares. @@ -315,7 +315,7 @@ In addition to the traditional scores, MODE output allows more information to be .. figure:: figure/plotting_fig4.jpg - Circle histograms showing object centroid angles and distances (see text for explanation). + Circle histograms showing object centroid angles and distances (see text for explanation). .. _TC-Stat-tool-example: @@ -324,7 +324,7 @@ TC-Stat Tool Example There is a basic R script located in the MET installation, *share/met/Rscripts/plot_tcmpr.R*. The usage statement with a short description of the options for *plot_tcmpr.R* can be obtained by typing: Rscript *plot_tcmpr.R* with no additional arguments. The only required argument is the **-lookin** source, which is the path to the TC-Pairs TCST output files. The R script reads directly from the TC-Pairs output, and calls TC-Stat directly for filter jobs specified in the *"-filter options"* argument. -In order to run this script, the MET_INSTALL_DIR environment variable must be set to the MET installation directory and the MET_BASE environment variable must be set to the *MET_INSTALL_DIR/share/met* directory. In addition, the Tc-Stat tool under *MET_INSTALL_DIR/bin* must be in your system path. +In order to run this script, the MET_INSTALL_DIR environment variable must be set to the MET installation directory and the MET_BASE environment variable must be set to the *MET_INSTALL_DIR/share/met* directory. In addition, the TC-Stat tool under *MET_INSTALL_DIR/bin* must be in your system path. The supplied R script can generate a number of different plot types including boxplots, mean, median, rank, and relative performance. Pairwise differences can be plotted for the boxplots, mean, and median. Normal confidence intervals are applied to all figures unless the no_ci option is set to TRUE. Below are two example plots generated from the tools. @@ -332,10 +332,10 @@ The supplied R script can generate a number of different plot types including bo .. figure:: figure/plotting_fig5.jpg - Example boxplot from plot_tcmpr.R. Track error distributions by lead time for three operational models GFNI, GHMI, HFWI. + Example boxplot from plot_tcmpr.R. Track error distributions by lead time for three operational models GFNI, GHMI, HWFI. .. _plotting_fig6: .. figure:: figure/plotting_fig6.jpg - Example mean intensity error with confidence intervals at 95% from plot_tcmpr.R. Raw intensity error by lead time for a homogeneous comparison of two operational models GHMI, HWFI. + Example mean intensity error with confidence intervals at 95% from plot_tcmpr.R. Raw intensity error by lead time for a homogeneous comparison of two operational models GHMI, HWFI. diff --git a/docs/Users_Guide/point-stat.rst b/docs/Users_Guide/point-stat.rst index 4442a6d565..753d174243 100644 --- a/docs/Users_Guide/point-stat.rst +++ b/docs/Users_Guide/point-stat.rst @@ -23,9 +23,9 @@ Interpolation/Matching Methods This section provides information about the various methods available in MET to match gridded model output to point observations. Matching in the vertical and horizontal are completed separately using different methods. -In the vertical, if forecasts and observations are at the same vertical level, then they are paired as-is. If any discrepancy exists between the vertical levels, then the forecasts are interpolated to the level of the observation. The vertical interpolation is done in the natural log of pressure coordinates, except for specific humidity, which is interpolated using the natural log of specific humidity in the natural log of pressure coordinates. Vertical interpolation for heights above ground are done linear in height coordinates. When forecasts are for the surface, no interpolation is done. They are matched to observations with message types that are mapped to "SURFACE" in the **message_type_group_map** configuration option. By default, the surface message types include ADPSFC, SFCSHP, and MSONET. The regular expression is applied to the message type list at the message_type_group_map. The derived message types from the time summary ("ADPSFC_MIN_hhmmss" and "ADPSFC_MAX_hhmmss") are accepted as "ADPSFC". +In the vertical, if forecasts and observations are at the same vertical level, then they are paired as-is. If any discrepancy exists between the vertical levels, then the forecasts are interpolated to the level of the observation. The vertical interpolation is done in the natural log of pressure coordinates, except for specific humidity, which is interpolated using the natural log of specific humidity in the natural log of pressure coordinates. Vertical interpolation for heights above ground is done linearly in height coordinates. When forecasts are for the surface, no interpolation is done. They are matched to observations with message types that are mapped to "SURFACE" in the **message_type_group_map** configuration option. By default, the surface message types include ADPSFC, SFCSHP, and MSONET. The regular expression is applied to the message type list at the message_type_group_map. The derived message types from the time summary ("ADPSFC_MIN_hhmmss" and "ADPSFC_MAX_hhmmss") are accepted as "ADPSFC". -To match forecasts and observations in the horizontal plane, the user can select from a number of methods described below. Many of these methods require the user to define the width of the forecast grid W, around each observation point P, that should be considered. In addition, the user can select the interpolation shape, either a SQUARE or a CIRCLE. For example, a square of width 2 defines the 2 x 2 set of grid points enclosing P, or simply the 4 grid points closest to P. A square of width of 3 defines a 3 x 3 square consisting of 9 grid points centered on the grid point closest to P. :numref:`point_stat_fig1` provides illustration. The point P denotes the observation location where the interpolated value is calculated. The interpolation width W, shown is five. +To match forecasts and observations in the horizontal plane, the user can select from a number of methods described below. Many of these methods require the user to define the width of the forecast grid W, around each observation point P, that should be considered. In addition, the user can select the interpolation shape, either a SQUARE or a CIRCLE. For example, a square of width 2 defines the 2 x 2 set of grid points enclosing P, or simply the 4 grid points closest to P. A square of width 3 defines a 3 x 3 square consisting of 9 grid points centered on the grid point closest to P. :numref:`point_stat_fig1` provides illustration. The point P denotes the observation location where the interpolated value is calculated. The interpolation width W, shown is five. This section describes the options for interpolation in the horizontal. @@ -33,13 +33,13 @@ This section describes the options for interpolation in the horizontal. .. figure:: figure/point_stat_fig1.png - Diagram illustrating matching and interpolation methods used in MET. See text for explanation. + Diagram illustrating matching and interpolation methods used in MET. See text for explanation. .. _point_stat_fig2: .. figure:: figure/point_stat_fig2.jpg - Illustration of some matching and interpolation methods used in MET. See text for explanation. + Illustration of some matching and interpolation methods used in MET. See text for explanation. ____________________ @@ -57,7 +57,7 @@ _____________________ **Gaussian** -The forecast value at P is a weighted sum of the values in the interpolation area. The weight given to each forecast point follows the Gaussian distribution with nearby points contributing more the far away points. The shape of the distribution is configured using sigma. +The forecast value at P is a weighted sum of the values in the interpolation area. The weight given to each forecast point follows the Gaussian distribution with nearby points contributing more than far away points. The shape of the distribution is configured using sigma. When used for regridding, with the **regrid** configuration option, or smoothing, with the **interp** configuration option in grid-to-grid comparisons, the Gaussian method is named **MAXGAUSS** and is implemented as a 2-step process. First, the data is regridded or smoothed using the maximum value interpolation method described below, where the **width** and **shape** define the interpolation area. Second, the Gaussian smoother, defined by the **gaussian_dx** and **gaussian_radius** configuration options, is applied. @@ -97,7 +97,7 @@ _____________________ To perform least squares interpolation of a gridded field at a location P, MET uses an **WxW** subgrid centered (as closely as possible) at P. :numref:`point_stat_fig1` shows the case where W = 5. -If we denote the horizontal coordinate in this subgrid by x, and vertical coordinate by y, then we can assign coordinates to the point P relative to this subgrid. These coordinates are chosen so that the center of the grid is. For example, in :numref:`point_stat_fig1`, P has coordinates (-0.4, 0.2). Since the grid is centered near P, the coordinates of P should always be at most 0.5 in absolute value. At each of the vertices of the grid (indicated by black dots in the figure), we have data values. We would like to use these values to interpolate a value at P. We do this using least squares. If we denote the interpolated value by z, then we fit an expression of the form :math:`z=\alpha (x) + \beta (y) + \gamma` over the subgrid. The values of :math:`\alpha, \beta, \gamma` are calculated from the data values at the vertices. Finally, the coordinates (**x,y**) of P are substituted into this expression to give z, our least squares interpolated data value at P. +If we denote the horizontal coordinate in this subgrid by x, and vertical coordinate by y, then we can assign coordinates to the point P relative to this subgrid. These coordinates are chosen so that the center of the grid is at the origin (0, 0). For example, in :numref:`point_stat_fig1`, P has coordinates (-0.4, 0.2). Since the grid is centered near P, the coordinates of P should always be at most 0.5 in absolute value. At each of the vertices of the grid (indicated by black dots in the figure), we have data values. We would like to use these values to interpolate a value at P. We do this using least squares. If we denote the interpolated value by z, then we fit an expression of the form :math:`z=\alpha (x) + \beta (y) + \gamma` over the subgrid. The values of :math:`\alpha, \beta, \gamma` are calculated from the data values at the vertices. Finally, the coordinates (**x,y**) of P are substituted into this expression to give z, our least squares interpolated data value at P. _______________________ @@ -123,12 +123,12 @@ Wind Rotation and Derivation ---------------------------- Numerical weather prediction model output often defines winds relative to the orientation of the model grid rather than true north-south and east-west directions on the earth. -However point observations typically define winds relative to true earth directions. Prior to comparing them, the model wind data must be rotated from grid-relative to +However point observations typically define winds relative to true earth directions. Prior to comparing them, the model wind data must be rotated from grid-relative to earth-relative. While the degree of grid-to-earth rotation varies based on the projection type and location, failing to rotate the winds can have a significant impact on the verification results. For simplicity, the MET library code attempts to rotate all wind components from being grid-relative to earth-relative regardless of whether they are being compared to point -observations with the Point-Stat tool or gridded analyses with the Grid-Stat tool. Running the MET tools at verbosity level 3 (-v 3) prints log messages to describing the wind +observations with the Point-Stat tool or gridded analyses with the Grid-Stat tool. Running the MET tools at verbosity level 3 (-v 3) prints log messages describing the wind rotation process. While the logic to determine whether input winds are grid-relative varies by file type, specifying the **is_grid_relative = FALSE;** configuration option manually overrides that logic and prevents winds from being rotated. @@ -138,11 +138,11 @@ required to rotate either component. When processing U-wind data, MET attempts t The configuration options for wind rotation and derivation in MET are described in section :numref:`config_wind_field_names`. When reading V-wind data to rotate U-wind or U-wind data to rotate V-wind, MET first searches using the same field name, but with both upper and lowercase U's and V's -swapped (e.g. for "U_PL" search for "V_PL"). If the result is unsuccesful, it searches other common field names specified by the **u_wind_field_name** and **v_wind_field_name** +swapped (e.g., for "U_PL" search for "V_PL"). If the result is unsuccessful, it searches other common field names specified by the **u_wind_field_name** and **v_wind_field_name** configuration options. If needed, users should set these configuration options to indicate how the U-wind and V-wind data should be paired. In addition to rotating winds, MET can also derive them. If U-wind and V-wind are present in the input file, request field names of **WDIR**, **WIND**, or **KENG** -to derive wind direction, wind speed, and kinetic engery from the U and V components, respectively. If wind speed and direction are present in the input file, request field +to derive wind direction, wind speed, and kinetic energy from the U and V components, respectively. If wind speed and direction are present in the input file, request field names of **UGRD** or **VGRD** for MET to derive the U and V components from them. Depending on the wind field naming conventions, the configuration options described in :numref:`config_wind_field_names` may be required to configure this derivation logic. @@ -155,13 +155,13 @@ The Point-Stat tool has been enhanced to include the High Resolution Assessment The HiRA framework provides a unique method for evaluating models in the neighborhood of point observations, allowing for some spatial and temporal uncertainty in the forecast and/or the observations. Additionally, the HiRA framework can be used to compare deterministic forecasts to ensemble forecasts. In MET, the neighborhood is a circle or square centered on the grid point closest to the observation location. An event is defined, then the proportion of points with events in the neighborhood is calculated. This proportion is treated as an ensemble probability, though it is likely to be uncalibrated. -:numref:`point_stat_fig3` shows a couple of examples of how the HiRA proportion is derived at a single model level using square neighborhoods. Events (in our case, model accretion values > 0) are separated from non-events (model accretion value = 0). Then, in each neighborhood, the total proportion of events is calculated. In the leftmost panel, four events exist in the 25 point neighborhood, making the HiRA proportion is 4/25 = 0.16. For the neighborhood of size 9 centered in that same panel, the HiRA proportion is 1/9. In the right panel, the size 25 neighborhood has HiRA proportion of 6/25, with the centered 9-point neighborhood having a HiRA value of 2/9. To extend this method into 3-dimensions, all layers within the user-defined layer are also included in the calculation of the proportion in the same manner. +:numref:`point_stat_fig3` shows a couple of examples of how the HiRA proportion is derived at a single model level using square neighborhoods. Events (in our case, model accretion values > 0) are separated from non-events (model accretion value = 0). Then, in each neighborhood, the total proportion of events is calculated. In the leftmost panel, four events exist in the 25 point neighborhood, making the HiRA proportion 4/25 = 0.16. For the neighborhood of size 9 centered in that same panel, the HiRA proportion is 1/9. In the right panel, the size 25 neighborhood has HiRA proportion of 6/25, with the centered 9-point neighborhood having a HiRA value of 2/9. To extend this method into 3-dimensions, all layers within the user-defined layer are also included in the calculation of the proportion in the same manner. .. _point_stat_fig3: .. figure:: figure/point_stat_fig3.png - Example showing how HiRA proportions are calculated. + Example showing how HiRA proportions are calculated. Often, the neighborhood size is chosen so that multiple models to be compared have approximately the same horizontal resolution. Then, standard metrics for probabilistic forecasts, such as Brier Score, can be used to compare those forecasts. HiRA was developed using surface observation stations so the neighborhood lies completely within the horizontal plane. With any type of upper air observation, the vertical neighborhood must also be defined. @@ -195,7 +195,7 @@ Measures for Continuous Variables For continuous variables, many verification measures are based on the forecast error (i.e., f - o). However, it also is of interest to investigate characteristics of the forecasts, and the observations, as well as their relationship. These concepts are consistent with the general framework for verification outlined by :ref:`Murphy and Winkler (1987) `. The statistics produced by MET for continuous forecasts represent this philosophy of verification, which focuses on a variety of aspects of performance rather than a single measure. See :numref:`Appendix C, Section %s ` for specific information. -A user may wish to eliminate certain values of the forecasts from the calculation of statistics, a process referred to here as``'conditional verification''. For example, a user may eliminate all temperatures above freezing and then calculate the error statistics only for those forecasts of below freezing temperatures. Another common example involves verification of wind forecasts. Since wind direction is indeterminate at very low wind speeds, the user may wish to set a minimum wind speed threshold prior to calculating error statistics for wind direction. The user may specify these thresholds in the configuration file to specify the conditional verification. Thresholds can be specified using the usual Fortran conventions (<, <=, ==, !-, >=, or >) followed by a numeric value. The threshold type may also be specified using two letter abbreviations (lt, le, eq, ne, ge, gt). Further, more complex thresholds can be achieved by defining multiple thresholds and using && or || to string together event definition logic. The forecast and observation threshold can be used together according to user preference by specifying one of: UNION, INTERSECTION, or SYMDIFF (symmetric difference). +A user may wish to eliminate certain values of the forecasts from the calculation of statistics, a process referred to here as "conditional verification". For example, a user may eliminate all temperatures above freezing and then calculate the error statistics only for those forecasts of below freezing temperatures. Another common example involves verification of wind forecasts. Since wind direction is indeterminate at very low wind speeds, the user may wish to set a minimum wind speed threshold prior to calculating error statistics for wind direction. The user may specify these thresholds in the configuration file to specify the conditional verification. Thresholds can be specified using the usual Fortran conventions (<, <=, ==, !=, >=, or >) followed by a numeric value. The threshold type may also be specified using two letter abbreviations (lt, le, eq, ne, ge, gt). Further, more complex thresholds can be achieved by defining multiple thresholds and using && or || to string together event definition logic. The forecast and observation threshold can be used together according to user preference by specifying one of: UNION, INTERSECTION, or SYMDIFF (symmetric difference). .. _PS_Probability: @@ -208,13 +208,13 @@ Probabilistic forecast values are assumed to have a range of either 0 to 1 or 0 MET supports multiple methods for defining probability bins: -1. As an explicit list of greater-than-or-equal-to-type thresholds whose values begin at 0.0, end at 1.0, and are monotonically increasing. For example :code:`>=0.00,>=0.25,>=0.50,>=0.75,>=1.00` defines 4 probability bins of equal width. Explicity listing the thresholds enables the definition of non-equal bin widths, such as :code:`>=0.00,>=0.50,>=0.75,>=1.00` with one bin of width 0.5 followed by two of width 0.25. +1. As an explicit list of greater-than-or-equal-to-type thresholds whose values begin at 0.0, end at 1.0, and are monotonically increasing. For example :code:`>=0.00,>=0.25,>=0.50,>=0.75,>=1.00` defines 4 probability bins of equal width. Explicitly listing the thresholds enables the definition of non-equal bin widths, such as :code:`>=0.00,>=0.50,>=0.75,>=1.00` with one bin of width 0.5 followed by two of width 0.25. 2. Since equal bin widths are commonly used, a shorthand notation of :code:`==0.25` is also supported. This defines the same 4 probability bins of equal width between 0 and 1, as shown in the example above. -3. As of MET version 12.0.0, an additional shorthand notation of :code:`==N`, where :code:`N` is an integer greater than 1, is also supported. With this notation, :code:`N` is interpreted as the number of ensemble members from which probabilities have been dervied. Often ensemble-derived probabilities are limited to N + 1 values, 0/N, 1/N, 2/N ... N/N. For :code:`==N`, MET defines thresholds to create N + 1 probability bins centered on each of the possible outcomes. For example, :code:`==4` expands to :code:`>=-0.125,>=0.125,>=0.375,>=0.625,>=0.875,>=1.125` which has 5 bins, each centered on 0.00, 0.25, 0.50, 0.75, and 1.00. Note that this convention results thresholds starting less than 0.0 and extending greater than 1.0. While this looks odd, it is necessary to create bins centered of the values of 0.0 and 1.0. +3. As of MET version 12.0.0, an additional shorthand notation of :code:`==N`, where :code:`N` is an integer greater than 1, is also supported. With this notation, :code:`N` is interpreted as the number of ensemble members from which probabilities have been derived. Often ensemble-derived probabilities are limited to N + 1 values, 0/N, 1/N, 2/N ... N/N. For :code:`==N`, MET defines thresholds to create N + 1 probability bins centered on each of the possible outcomes. For example, :code:`==4` expands to :code:`>=-0.125,>=0.125,>=0.375,>=0.625,>=0.875,>=1.125` which has 5 bins, each centered on 0.00, 0.25, 0.50, 0.75, and 1.00. Note that this convention results in thresholds starting less than 0.0 and extending greater than 1.0. While this looks odd, it is necessary to create bins centered on the values of 0.0 and 1.0. -When computing probabilistic statistics, MET first bins the forecast probabilities into an Nx2 probabilistic contingency table. Probabilistic statistics are derived from the Nx2 contingency table rather than the raw probabilities. When doing so, the value of the mid-point is used for all values falling in that bin. Because of this, the choice of probability bins impacts the statistics. For a well-calibrated, smooth distribution of probabilities, the impact of binning is relatively minor. However the impact may be larger for less continuous probabilities. In particular, when evaluating ensemble-derived probability values, users are encouarged to define probability bins with the :code:`==N` option to create bins centered on the possible ensemble-derived probability outcomes. +When computing probabilistic statistics, MET first bins the forecast probabilities into an Nx2 probabilistic contingency table. Probabilistic statistics are derived from the Nx2 contingency table rather than the raw probabilities. When doing so, the value of the mid-point is used for all values falling in that bin. Because of this, the choice of probability bins impacts the statistics. For a well-calibrated, smooth distribution of probabilities, the impact of binning is relatively minor. However the impact may be larger for less continuous probabilities. In particular, when evaluating ensemble-derived probability values, users are encouraged to define probability bins with the :code:`==N` option to create bins centered on the possible ensemble-derived probability outcomes. When the "prob" entry is set as a dictionary to define the field of interest, setting "prob_as_scalar = TRUE" indicates that this data should be processed as regular scalars rather than probabilities. For example, this option can be used to compute traditional 2x2 contingency tables and neighborhood verification statistics for probability data. It can also be used to compare two probability fields directly. @@ -223,7 +223,7 @@ When the "prob" entry is set as a dictionary to define the field of interest, se Measures for Comparison Against Climatology ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ -For each of the types of statistics mentioned above (categorical, continuous, and probabilistic), it is possible to calculate measures of skill relative to climatology. MET will accept a climatology file provided by the user, and will evaluate it as a reference forecast. Further, anomalies, i.e. departures from average conditions, can be calculated. As with all other statistics, the available measures will depend on the nature of the forecast. Common statistics that use a climatological reference include: the mean squared error skill score (MSESS), the Anomaly Correlation (ANOM_CORR and ANOM_CORR_UNCNTR), scalar and vector anomalies (SAL1L2 and VAL1L2), continuous ranked probability skill score (CRPSS and CRPSS_EMP), Brier Skill Score (BSS) (:ref:`Wilks, 2011 `; :ref:`Mason, 2004 `). +For each of the types of statistics mentioned above (categorical, continuous, and probabilistic), it is possible to calculate measures of skill relative to climatology. MET will accept a climatology file provided by the user, and will evaluate it as a reference forecast. Further, anomalies, i.e., departures from average conditions, can be calculated. As with all other statistics, the available measures will depend on the nature of the forecast. Common statistics that use a climatological reference include: the mean squared error skill score (MSESS), the Anomaly Correlation (ANOM_CORR and ANOM_CORR_UNCNTR), scalar and vector anomalies (SAL1L2 and VAL1L2), continuous ranked probability skill score (CRPSS and CRPSS_EMP), Brier Skill Score (BSS) (:ref:`Wilks, 2011 `; :ref:`Mason, 2004 `). Often, the sample climatology is used as a reference by a skill score. The sample climatology is the average over all included observations and may be transparent to the user. This is the case in most categorical skill scores. The sample climatology will probably prove more difficult to improve upon than a long term climatology, since it will be from the same locations and time periods as the forecasts. This may mask legitimate forecast skill. However, a more general climatology, perhaps covering many years, is often easier to improve upon and is less likely to mask real forecast skill. @@ -238,7 +238,7 @@ For continuous fields (e.g., temperature), it is possible to estimate confidence For the measures relating the two fields (i.e., mean error, correlation and standard deviation of the errors), confidence intervals are based on either the joint distributions of the two fields (e.g., with correlation) or on a function of the two fields. For the correlation, the underlying assumption is that the two fields follow a bivariate normal distribution. In the case of the mean error and the standard deviation of the mean error, the assumption is that the errors are normally distributed, which for continuous variables, is usually a reasonable assumption, even for the standard deviation of the errors. -Bootstrap confidence intervals for any verification statistic are available in MET. Bootstrapping is a nonparametric statistical method for estimating parameters and uncertainty information. The idea is to obtain a sample of the verification statistic(s) of interest (e.g., bias, ETS, etc.) so that inferences can be made from this sample. The assumption is that the original sample of matched forecast-observation pairs is representative of the population. Several replicated samples are taken with replacement from this set of forecast-observation pairs of variables (e.g., precipitation, temperature, etc.), and the statistic(s) are calculated for each replicate. That is, given a set of n forecast-observation pairs, we draw values at random from these pairs, allowing the same pair to be drawn more than once, and the statistic(s) is (are) calculated for each replicated sample. This yields a sample of the statistic(s) based solely on the data without making any assumptions about the underlying distribution of the sample. It should be noted, however, that if the observed sample of matched pairs is dependent, then this dependence should be taken into account somehow. Currently, the confidence interval methods in MET do not take into account dependence, but future releases will support a robust method allowing for dependence in the original sample. More detailed information about the bootstrap algorithm is found in the :numref:`Appendix D, Section %s `. Note that MET writes temporary files whenever bootstrap confidence intervals are computed, as described in :numref:`Contributor's Guide Section %s `. +Bootstrap confidence intervals for any verification statistic are available in MET. Bootstrapping is a nonparametric statistical method for estimating parameters and uncertainty information. The idea is to obtain a sample of the verification statistic(s) of interest (e.g., bias, ETS, etc.) so that inferences can be made from this sample. The assumption is that the original sample of matched forecast-observation pairs is representative of the population. Several replicated samples are taken with replacement from this set of forecast-observation pairs of variables (e.g., precipitation, temperature, etc.), and the statistic(s) are calculated for each replicate. That is, given a set of n forecast-observation pairs, we draw values at random from these pairs, allowing the same pair to be drawn more than once, and the statistic(s) is (are) calculated for each replicated sample. This yields a sample of the statistic(s) based solely on the data without making any assumptions about the underlying distribution of the sample. It should be noted, however, that if the observed sample of matched pairs is dependent, then this dependence should be taken into account somehow. Currently, the confidence interval methods in MET do not take into account dependence, but future releases will support a robust method allowing for dependence in the original sample. More detailed information about the bootstrap algorithm is found in :numref:`Appendix D, Section %s `. Note that MET writes temporary files whenever bootstrap confidence intervals are computed, as described in :numref:`Contributor's Guide Section %s `. Confidence intervals can be calculated from the sample of verification statistics obtained through the bootstrap algorithm. The most intuitive method is to simply take the appropriate quantiles of the sample of statistic(s). For example, if one wants a 95% CI, then one would take the 2.5 and 97.5 percentiles of the resulting sample. This method is called the percentile method, and has some nice properties. However, if the original sample is biased and/or has non-constant variance, then it is well known that this interval is too optimistic. The most robust, accurate, and well-behaved way to obtain accurate CIs from bootstrapping is to use the bias corrected and adjusted percentile method (or BCa). If there is no bias, and the variance is constant, then this method will yield the usual percentile interval. The only drawback to the approach is that it is computationally intensive. Therefore, both the percentile and BCa methods are available in MET, with the considerably more efficient percentile method being the default. @@ -371,7 +371,7 @@ point_stat Configuration File The default configuration file for the Point-Stat tool named **PointStatConfig_default** can be found in the installed *share/met/config* directory. Another version is located in *scripts/config*. We encourage users to make a copy of these files prior to modifying their contents. The contents of the configuration file are described in the subsections below. -Note that environment variables may be used when editing configuration files, as described in the :numref:`config_env_vars`. +Note that environment variables may be used when editing configuration files, as described in :numref:`config_env_vars`. ________________________ @@ -421,9 +421,9 @@ _________________________ Setting up the **fcst** and **obs** dictionaries of the configuration file is described in :numref:`config_options`. The following are some special considerations for the Point-Stat tool. -The **obs** dictionary looks very similar to the **fcst** dictionary. When the forecast and observation variables follow the same naming convention, one can easily copy over the forecast settings to the observation dictionary using **obs = fcst;**. However when verifying forecast data in NetCDF format or verifying against not-standard observation variables, users will need to specify the **fcst** and **obs** dictionaries separately. The number of fields specified in the **fcst** and **obs** dictionaries must match. +The **obs** dictionary looks very similar to the **fcst** dictionary. When the forecast and observation variables follow the same naming convention, one can easily copy over the forecast settings to the observation dictionary using **obs = fcst;**. However when verifying forecast data in NetCDF format or verifying against non-standard observation variables, users will need to specify the **fcst** and **obs** dictionaries separately. The number of fields specified in the **fcst** and **obs** dictionaries must match. -The **message_type** entry, defined in the **obs** dictionary, contains a comma-separated list of the message types to use for verification. At least one entry must be provided. The Point-Stat tool performs verification using observations for one message type at a time. See `Table 1.a Current Table A Entries in PREPBUFR mnemonic table `_ for a list of the possible types. If using **obs = fcst;**, it can be defined in the forecast dictionary and the copied into the observation dictionary. +The **message_type** entry, defined in the **obs** dictionary, contains a comma-separated list of the message types to use for verification. At least one entry must be provided. The Point-Stat tool performs verification using observations for one message type at a time. See `Table 1.a Current Table A Entries in PREPBUFR mnemonic table `_ for a list of the possible types. If using **obs = fcst;**, it can be defined in the forecast dictionary and then copied into the observation dictionary. ________________________ @@ -438,7 +438,7 @@ ________________________ prob_cat_thresh = []; } -The **hira** dictionary that is very similar to the **interp** and **nbrhd** entries. It specifies information for applying the High Resolution Assessment (HiRA) verification logic described in section :numref:`PS_HiRA_framework`. The **flag** entry is a boolean which toggles HiRA on (**TRUE**) and off (**FALSE**). The **width** and **shape** entries define the neighborhood size and shape, respectively. Since HiRA applies to point observations, the width may be even or odd. The **vld_thresh** entry is the required ratio of valid data within the neighborhood to compute an output value. The **cov_thresh** entry is an array of probabilistic thresholds used to populate the Nx2 probabilistic contingency table written to the PCT output line and used for computing probabilistic statistics. The **prob_cat_thresh** entry defines the thresholds to be used in computing the ranked probability score in the RPS output line type. If left empty but climatology data is provided, the **climo_cdf** thresholds will be used instead of **prob_cat_thresh**. +The **hira** dictionary is very similar to the **interp** and **nbrhd** entries. It specifies information for applying the High Resolution Assessment (HiRA) verification logic described in section :numref:`PS_HiRA_framework`. The **flag** entry is a boolean which toggles HiRA on (**TRUE**) and off (**FALSE**). The **width** and **shape** entries define the neighborhood size and shape, respectively. Since HiRA applies to point observations, the width may be even or odd. The **vld_thresh** entry is the required ratio of valid data within the neighborhood to compute an output value. The **cov_thresh** entry is an array of probabilistic thresholds used to populate the Nx2 probabilistic contingency table written to the PCT output line and used for computing probabilistic statistics. The **prob_cat_thresh** entry defines the thresholds to be used in computing the ranked probability score in the RPS output line type. If left empty but climatology data is provided, the **climo_cdf** thresholds will be used instead of **prob_cat_thresh**. ________________________ @@ -718,7 +718,7 @@ The first set of header columns are common to all of the output files generated - Double .. role:: raw-html(raw) - :format: html + :format: html .. _table_PS_format_info_CTS: @@ -795,23 +795,23 @@ The first set of header columns are common to all of the output files generated - Logarithm of the Odds Ratio including normal and bootstrap upper and lower confidence limits - Double * - 90-94 - - ORSS, :raw-html:`
` ORSS _NCL, :raw-html:`
` ORSS _NCU, :raw-html:`
` ORSS _BCL, :raw-html:`
` ORSS _BCU + - ORSS, :raw-html:`
` ORSS_NCL, :raw-html:`
` ORSS_NCU, :raw-html:`
` ORSS_BCL, :raw-html:`
` ORSS_BCU - Odds Ratio Skill Score including normal and bootstrap upper and lower confidence limits - Double * - 95-99 - - EDS, :raw-html:`
` EDS _NCL, :raw-html:`
` EDS _NCU, :raw-html:`
` EDS _BCL, :raw-html:`
` EDS _BCU + - EDS, :raw-html:`
` EDS_NCL, :raw-html:`
` EDS_NCU, :raw-html:`
` EDS_BCL, :raw-html:`
` EDS_BCU - Extreme Dependency Score including normal and bootstrap upper and lower confidence limits - Double * - 100-104 - - SEDS, :raw-html:`
` SEDS _NCL, :raw-html:`
` SEDS _NCU, :raw-html:`
` SEDS _BCL, :raw-html:`
` SEDS _BCU + - SEDS, :raw-html:`
` SEDS_NCL, :raw-html:`
` SEDS_NCU, :raw-html:`
` SEDS_BCL, :raw-html:`
` SEDS_BCU - Symmetric Extreme Dependency Score including normal and bootstrap upper and lower confidence limits - Double * - 105-109 - - EDI, :raw-html:`
` EDI _NCL, :raw-html:`
` EDI _NCU, :raw-html:`
` EDI _BCL, :raw-html:`
` EDI _BCU + - EDI, :raw-html:`
` EDI_NCL, :raw-html:`
` EDI_NCU, :raw-html:`
` EDI_BCL, :raw-html:`
` EDI_BCU - Extreme Dependency Index including normal and bootstrap upper and lower confidence limits - Double * - 111-113 - - SEDI, :raw-html:`
` SEDI _NCL, :raw-html:`
` SEDI _NCU, :raw-html:`
` SEDI _BCL, :raw-html:`
` SEDI _BCU + - SEDI, :raw-html:`
` SEDI_NCL, :raw-html:`
` SEDI_NCU, :raw-html:`
` SEDI_BCL, :raw-html:`
` SEDI_BCU - Symmetric Extremal Dependency Index including normal and bootstrap upper and lower confidence limits - Double * - 115-117 @@ -829,7 +829,7 @@ The first set of header columns are common to all of the output files generated .. role:: raw-html(raw) - :format: html + :format: html .. _table_PS_format_info_CNT: @@ -922,7 +922,7 @@ The first set of header columns are common to all of the output files generated - 10th, 25th, 50th, 75th, and 90th percentiles of the error including bootstrap upper and lower confidence limits - Double * - 96-98 - - EIQR, :raw-html:`
` IQR _BCL, :raw-html:`
` IQR _BCU + - EIQR, :raw-html:`
` EIQR_BCL, :raw-html:`
` EIQR_BCU - The Interquartile Range of the error including bootstrap upper and lower confidence limits - Double * - 99-101 @@ -992,7 +992,7 @@ The first set of header columns are common to all of the output files generated .. role:: raw-html(raw) - :format: html + :format: html .. _table_PS_format_info_MCTS: @@ -1082,7 +1082,7 @@ The first set of header columns are common to all of the output files generated .. role:: raw-html(raw) - :format: html + :format: html .. _table_PS_format_info_PSTD: @@ -1159,7 +1159,7 @@ The first set of header columns are common to all of the output files generated - Data Type * - 24 - PJC - - Probabilistic Joint/Continuous line type + - Probabilistic Joint and Conditional factorization line type - String * - 25 - TOTAL @@ -1422,7 +1422,7 @@ The first set of header columns are common to all of the output files generated - Double * - 35 - TOTAL_DIR - - Total number of matched pairs for which both the forecast and observation wind directions are well-defined (i.e. non-zero vectors) + - Total number of matched pairs for which both the forecast and observation wind directions are well-defined (i.e., non-zero vectors) - Double * - 36 - DIR_ME @@ -1493,7 +1493,7 @@ The first set of header columns are common to all of the output files generated - Double * - 35 - TOTAL_DIR - - Total number of matched pairs for which the forecast, observation, forecast climatology, and observation climatology wind directions are well-defined (i.e. non-zero vectors) + - Total number of matched pairs for which the forecast, observation, forecast climatology, and observation climatology wind directions are well-defined (i.e., non-zero vectors) - Double * - 36 - DIRA_ME @@ -1580,7 +1580,7 @@ The first set of header columns are common to all of the output files generated - Double * - 65-67 - VDIFF_DIR, :raw-html:`
` VDIFF_DIR_BCL, :raw-html:`
` VDIFF_DIR_BCU - - Direction of the vector difference between the average forecast and average wind vectors including bootstrap upper and lower confidence limits + - Direction of the vector difference between the average forecast and average observed wind vectors including bootstrap upper and lower confidence limits - Double * - 68-70 - SPEED_ERR, :raw-html:`
` SPEED_ERR_BCL, :raw-html:`
` SPEED_ERR_BCU @@ -1596,7 +1596,7 @@ The first set of header columns are common to all of the output files generated - Double * - 77-79 - DIR_ABSERR, :raw-html:`
` DIR_ABSERR_BCL, :raw-html:`
` DIR_ABSERR_BCU - - Absolute value of DIR_ABSERR including bootstrap upper and lower confidence limits + - Absolute value of DIR_ERR including bootstrap upper and lower confidence limits - Double * - 80-84 - ANOM_CORR, :raw-html:`
` ANOM_CORR_NCL, :raw-html:`
` ANOM_CORR_NCU, :raw-html:`
` ANOM_CORR_BCL, :raw-html:`
` ANOM_CORR_BCU @@ -1608,7 +1608,7 @@ The first set of header columns are common to all of the output files generated - Double * - 88 - TOTAL_DIR - - Total number of matched pairs for which both the forecast and observation wind directions are well-defined (i.e. non-zero vectors) + - Total number of matched pairs for which both the forecast and observation wind directions are well-defined (i.e., non-zero vectors) - Double * - 89-91 - DIR_ME, :raw-html:`
` DIR_ME_BCL, :raw-html:`
` DIR_ME_BCU diff --git a/docs/Users_Guide/reformat_grid.rst b/docs/Users_Guide/reformat_grid.rst index dd502a997d..abe9340e36 100644 --- a/docs/Users_Guide/reformat_grid.rst +++ b/docs/Users_Guide/reformat_grid.rst @@ -4,7 +4,7 @@ Re-Formatting of Gridded Fields ******************************* -Several MET tools exist for the purpose of reformatting gridded fields, and they are described in this section. These tools are represented by the reformatting column of MET flowchart depicted in :numref:`overview-figure`. +Several MET tools exist for the purpose of reformatting gridded fields, and they are described in this section. These tools are represented by the reformatting column of the MET flowchart depicted in :numref:`overview-figure`. Pcp-Combine Tool ================ @@ -21,7 +21,7 @@ The Pcp-Combine tool supports four types of commands ("sum", "add", "subtract", 4. The "derive" command reads the requested data from the input data files and computes the requested summary fields. -By default, the Pcp-Combine tool processes data for **APCP**, the GRIB string for accumulated precipitation. When requesting data using time strings (i.e. [HH]MMSS), Pcp-Combine searches for accumulated precipitation for that accumulation interval. Alternatively, use the "-field" option to process fields other than **APCP** or for non-GRIB files. The "-field" option may be used multiple times to process multiple fields in a single run. Since the Pcp-Combine tool does not support automated regridding, all input data must be on the same grid. In general the input files should have the same initialization time unless the user has indicated that it should ignore the initialization time for the "sum" command. The "subtract" command produces a warning when the input initialization times differ or the subtraction results in a negative accumulation interval. +By default, the Pcp-Combine tool processes data for **APCP**, the GRIB string for accumulated precipitation. When requesting data using time strings (i.e., [HH]MMSS), Pcp-Combine searches for accumulated precipitation for that accumulation interval. Alternatively, use the "-field" option to process fields other than **APCP** or for non-GRIB files. The "-field" option may be used multiple times to process multiple fields in a single run. Since the Pcp-Combine tool does not support automated regridding, all input data must be on the same grid. In general the input files should have the same initialization time unless the user has indicated that it should ignore the initialization time for the "sum" command. The "subtract" command produces a warning when the input initialization times differ or the subtraction results in a negative accumulation interval. pcp_combine Usage ----------------- @@ -76,7 +76,7 @@ Required Arguments for the pcp_combine Optional Arguments for pcp_combine ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ -3. The **-field string** option defines the data to be extracted from the input files. Use this option when processing fields other than **APCP** or non-GRIB files. It can be used multiple times and output will be created for each. In general, the field string should include the **name** and **level** of the requested data and be enclosed in single quotes. It is processed as an inline configuration file and may also include data filtering, censoring, and conversion options. For example, use **-field ‘name=”ACPCP”; level=”A6”; convert(x)=x/25.4;’** to read 6-hourly accumulated convective precipitation from a GRIB file and convert from millimeters to inches. +3. The **-field string** option defines the data to be extracted from the input files. Use this option when processing fields other than **APCP** or non-GRIB files. It can be used multiple times and output will be created for each. In general, the field string should include the **name** and **level** of the requested data and be enclosed in single quotes. It is processed as an inline configuration file and may also include data filtering, censoring, and conversion options. For example, use **-field 'name="ACPCP"; level="A6"; convert(x)=x/25.4;'** to read 6-hourly accumulated convective precipitation from a GRIB file and convert from millimeters to inches. 4. The **-name list** option is a comma-separated list of output variable names which override the default choices. If specified, the number of names must match the number of variables to be written to the output file. @@ -116,9 +116,9 @@ Required Arguments for the pcp_combine Derive Command Input Files for pcp_combine Add, Subtract, and Derive Commands ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ -The input files for the add, subtract, and derive command can be specified in one of 3 ways: +The input files for the add, subtract, and derive commands can be specified in one of 3 ways: -1. Use **file_1 config_str_1 ... file_n config_str_n** to specify the full path to each input file followed by a description of the data to be read from it. The **config_str_i** argument describing the data can be a set to a time string in HH[MMSS] format for accumulated precipitation or a full configuration string. For example, use **'name="TMP"; level="P500";'** to process temperature at 500mb. +1. Use **file_1 config_str_1 ... file_n config_str_n** to specify the full path to each input file followed by a description of the data to be read from it. The **config_str_i** argument describing the data can be set to a time string in HH[MMSS] format for accumulated precipitation or a full configuration string. For example, use **'name="TMP"; level="P500";'** to process temperature at 500mb. 2. Use **file_1 ... file_n** to specify the list of input files to be processed on the command line. Rather than specifying a separate configuration string for each input file, the "-field" command line option is required to specify the data to be processed. @@ -186,9 +186,9 @@ Each NetCDF file generated by the Pcp-Combine tool contains the dimensions and v * - NetCDF dimension - Description * - lat - - Dimension of the latitude (i.e. Number of grid points in the North-South direction) + - Dimension of the latitude (i.e., Number of grid points in the North-South direction) * - lon - - Dimension of the longitude (i.e. Number of grid points in the East-West direction) + - Dimension of the longitude (i.e., Number of grid points in the East-West direction) .. list-table:: NetCDF variables for pcp_combine output. @@ -209,7 +209,7 @@ Each NetCDF file generated by the Pcp-Combine tool contains the dimensions and v - Double * - Name and level of the requested data or value of the -name option. - lat, lon - - Data value (i.e. accumulated precipitation) for each point in the grid. The name of the variable describes the name and level and any derivation logic that was applied. + - Data value (i.e., accumulated precipitation) for each point in the grid. The name of the variable describes the name and level and any derivation logic that was applied. - Double .. _regrid-data-plane: @@ -217,7 +217,7 @@ Each NetCDF file generated by the Pcp-Combine tool contains the dimensions and v Regrid-Data-Plane Tool ====================== -This section contains a description of running the Regrid-Data-Plane tool. This tool may be run to read data from any gridded file MET supports, interpolate to a user-specified grid, and writes the field(s) out in NetCDF format. The user may specify the method of interpolation used for regridding as well as which fields to regrid. This tool is particularly useful when dealing with GRIB2 and NetCDF input files that need to be regridded. For GRIB1 files, it has also been tested for compatibility with the copygb regridding utility mentioned in :numref:`suggested_external_utiliites`. +This section contains a description of running the Regrid-Data-Plane tool. This tool may be run to read data from any gridded file MET supports, interpolate to a user-specified grid, and write the field(s) out in NetCDF format. The user may specify the method of interpolation used for regridding as well as which fields to regrid. This tool is particularly useful when dealing with GRIB2 and NetCDF input files that need to be regridded. For GRIB1 files, it has also been tested for compatibility with the copygb regridding utility mentioned in :numref:`suggested_external_utilities`. regrid_data_plane Usage ----------------------- @@ -284,7 +284,7 @@ For more details on setting the **to_grid, -method, -width,** and **-vld_thresh* input.grb \ togrid.grb \ regridded.nc \ - -field 'name="APCP"; level="A6";' + -field 'name="APCP"; level="A6";' \ -field 'name="TMP"; level="Z2";' \ -field 'name="UGRD"; level="Z10";' \ -field 'name="VGRD"; level="Z10";' \ @@ -318,7 +318,7 @@ The usage statement for the shift_data_plane utility is shown below: -to lat lon [-method type] [-width n] - [-shape SHAPE] + [-shape SHAPE] [-log file] [-v level] [-compress level] @@ -365,7 +365,7 @@ For more details on setting the **-method** and **-width** options, see the **re -to 40.1717 -105.1092 \ -v 2 -In this example, the Shift-Data-Plane tool reads 12-hour accumulated precipitation from the **nam.grb** file, applies a rigid shift defined by (38.6272, -90.1978) to (40.1717, -105.1092) and writes the output in NetCDF format to a file named **nam_shift_APCP_12.nc**. These **-from** and **-to** locations result in a grid shift of -108.30 units in the x-direction and 16.67 units in the y-direction. +In this example, the Shift-Data-Plane tool reads 12-hour accumulated precipitation from the **nam.grib** file, applies a rigid shift defined by (38.6272, -90.1978) to (40.1717, -105.1092) and writes the output in NetCDF format to a file named **nam_shift_APCP_12.nc**. These **-from** and **-to** locations result in a grid shift of -108.30 units in the x-direction and 16.67 units in the y-direction. MODIS regrid Tool ================= @@ -435,7 +435,7 @@ In this example, the Modis-Regrid tool will process the Cloud_Fraction field fro .. figure:: figure/reformat_grid_fig1.png - Example plot showing surface temperature from a MODIS file. + Example plot showing surface temperature from a MODIS file. WWMCA Tool Documentation ======================== @@ -458,7 +458,7 @@ The usage statement for the WWMCA-Plot tool is shown below: [-v level] wwmca_cloud_pct_file_list -wmmca_plot has some required arguments and can also take optional ones. +wwmca_plot has some required arguments and can also take optional ones. Required Arguments for wwmca_plot ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ @@ -478,7 +478,7 @@ Optional Arguments for wwmca_plot .. figure:: figure/reformat_grid_fig2.png - Example output of WWMCA-Plot tool. + Example output of WWMCA-Plot tool. wwmca_regrid Usage ------------------ @@ -496,12 +496,12 @@ The usage statement for the WWMCA-Regrid tool is shown below: [-v level] [-compress level] -wmmca_regrid has some required arguments and can also take optional ones. +wwmca_regrid has some required arguments and can also take optional ones. Required Arguments for wwmca_regrid ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ -1. The **-out filename** argument specifies the name of the output netCDF file. +1. The **-out filename** argument specifies the name of the output NetCDF file. 2. The **-config filename** argument indicates the name of the configuration file to be used. The contents of the configuration file are discussed below. @@ -525,7 +525,7 @@ wwmca_regrid Configuration File The default configuration file for the WWMCA-Regrid tool named **WWMCARegridConfig_default** can be found in the installed *share/met/config* directory. We encourage users to make a copy of this file prior to modifying its contents. The contents of the configuration file are described in the subsections below. -Note that environment variables may be used when editing configuration files, as described in the :numref:`config_env_vars`. +Note that environment variables may be used when editing configuration files, as described in :numref:`config_env_vars`. ____________________________ @@ -544,7 +544,7 @@ ____________________________ long_name = "cloud cover percent"; level = "SFC"; -The settings listed above are strings which control the output netCDF variable name and specify attributes for that variable. +The settings listed above are strings which control the output NetCDF variable name and specify attributes for that variable. ___________________________ diff --git a/docs/Users_Guide/reformat_point.rst b/docs/Users_Guide/reformat_point.rst index 1b7e816dfe..76150498ca 100644 --- a/docs/Users_Guide/reformat_point.rst +++ b/docs/Users_Guide/reformat_point.rst @@ -71,9 +71,9 @@ An example of the pb2nc calling sequence is shown below: .. code-block:: none - pb2nc sample_pb.blk \ - sample_pb.nc \ - PB2NCConfig + pb2nc sample_pb.blk \ + sample_pb.nc \ + PB2NCConfig In this example, the PB2NC tool will process the input **sample_pb.blk** file applying the configuration specified in the **PB2NCConfig** file and write the output to a file named **sample_pb.nc**. @@ -84,16 +84,16 @@ pb2nc Configuration File The default configuration file for the PB2NC tool named **PB2NCConfig_default** can be found in the installed *share/met/config* directory. The version used for the installation test cases is available in *scripts/config*. It is recommended that users make a copy of configuration files prior to modifying their contents. -Note that environment variables may be used when editing configuration files, as described in the :numref:`config_env_vars`. +Note that environment variables may be used when editing configuration files, as described in :numref:`config_env_vars`. ____________________ .. code-block:: none - obs_window = { beg = -5400; end = 5400; } - mask = { grid = ""; poly = ""; } - tmp_dir = "/tmp"; - version = "VN.N"; + obs_window = { beg = -5400; end = 5400; } + mask = { grid = ""; poly = ""; } + tmp_dir = "/tmp"; + version = "VN.N"; The configuration options listed above are common to many MET tools and are described in :numref:`config_options`. The use of temporary files in PB2NC is described in :numref:`Contributor's Guide Section %s `. @@ -102,7 +102,7 @@ _____________________ .. code-block:: none - message_type = []; + message_type = []; Each PrepBUFR message is tagged with one of eighteen message types as listed in the :numref:`config_options` file. The **message_type** refers to the type of observation from which the observation value (or 'report') was derived. The user may specify a comma-separated list of message types to be retained. Providing an empty list indicates that all message types should be retained. @@ -110,7 +110,7 @@ _____________________ .. code-block:: none - message_type_map = [ { key = "AIRCAR"; val = "AIRCAR_PROFILES"; } ]; + message_type_map = [ { key = "AIRCAR"; val = "AIRCAR_PROFILES"; } ]; The **message_type_map** entry is an array of dictionaries, each containing a **key** string and **val** string. This defines a mapping of input PrepBUFR message types to output message types. This provides a method for renaming input PrepBUFR message types. @@ -118,12 +118,12 @@ _____________________ .. code-block:: none - message_type_group_map = [ - { key = "SURFACE"; val = "ADPSFC,SFCSHP,MSONET"; }, - { key = "ANYAIR"; val = "AIRCAR,AIRCFT"; }, - { key = "ANYSFC"; val = "ADPSFC,SFCSHP,ADPUPA,PROFLR,MSONET"; }, - { key = "ONLYSF"; val = "ADPSFC,SFCSHP"; } - ]; + message_type_group_map = [ + { key = "SURFACE"; val = "ADPSFC,SFCSHP,MSONET"; }, + { key = "ANYAIR"; val = "AIRCAR,AIRCFT"; }, + { key = "ANYSFC"; val = "ADPSFC,SFCSHP,ADPUPA,PROFLR,MSONET"; }, + { key = "ONLYSF"; val = "ADPSFC,SFCSHP"; } + ]; The **message_type_group_map** entry is an array of dictionaries, each containing a **key** string and **val** string. This defines a mapping of message type group names to a comma-separated list of values. This map is defined in the config files for PB2NC, Point-Stat, or Ensemble-Stat. Modify this map to define sets of message types that should be processed together as a group. The **SURFACE** entry must be present to define message types for which surface verification logic should be applied. @@ -131,7 +131,7 @@ _____________________ .. code-block:: none - station_id = []; + station_id = []; Each PrepBUFR message has a station identification string associated with it. The user may specify a comma-separated list of station IDs to be retained. Providing an empty list indicates that messages from all station IDs will be retained. It can be a file name containing a list of stations. @@ -139,7 +139,7 @@ _____________________ .. code-block:: none - elevation_range = { beg = -1000; end = 100000; } + elevation_range = { beg = -1000; end = 100000; } The **beg** and **end** variables are used to stratify the elevation (in meters) of the observations to be retained. The range shown above is set to -1000 to 100000 meters, which essentially retains every observation. @@ -147,9 +147,9 @@ _____________________ .. code-block:: none - pb_report_type = []; - in_report_type = []; - instrument_type = []; + pb_report_type = []; + in_report_type = []; + instrument_type = []; The **pb_report_type, in_report_type**, and **instrument_type** variables are used to specify comma-separated lists of PrepBUFR report types, input report types, and instrument types to be retained, respectively. If left empty, all PrepBUFR report types, input report types, and instrument types will be retained. See the following for more details: @@ -161,8 +161,8 @@ _____________________ .. code-block:: none - level_range = { beg = 1; end = 255; } - level_category = []; + level_range = { beg = 1; end = 255; } + level_category = []; The **beg** and **end** variables are used to stratify the model level of observations to be retained. The range shown above is 1 to 255. @@ -173,27 +173,27 @@ The **level_category** variable is used to specify a comma-separated list of Pre .. _table_reformat-point_pb2nc_level_category: .. list-table:: Values for the level_category option. - :widths: auto - :header-rows: 1 - - * - Level category value - - Description - * - 0 - - Surface level - * - 1 - - Mandatory level - * - 2 - - Significant temperature level - * - 3 - - Winds-by-pressure level - * - 4 - - Winds-by-height level - * - 5 - - Tropopause level - * - 6 - - Reports on a single level - * - 7 - - Auxiliary levels generated via interpolation from spanning levels + :widths: auto + :header-rows: 1 + + * - Level category value + - Description + * - 0 + - Surface level + * - 1 + - Mandatory level + * - 2 + - Significant temperature level + * - 3 + - Winds-by-pressure level + * - 4 + - Winds-by-height level + * - 5 + - Tropopause level + * - 6 + - Reports on a single level + * - 7 + - Auxiliary levels generated via interpolation from spanning levels _____________________ @@ -229,73 +229,73 @@ _____________________ .. code-block:: none - obs_bufr_map = [ - { key = 'POB'; val = 'PRES'; }, - { key = 'QOB'; val = 'SPFH'; }, - { key = 'TOB'; val = 'TMP'; }, - { key = 'ZOB'; val = 'HGT'; }, - { key = 'UOB'; val = 'UGRD'; }, - { key = 'VOB'; val = 'VGRD'; }, - { key = 'D_DPT'; val = 'DPT'; }, - { key = 'D_WDIR'; val = 'WDIR'; }, - { key = 'D_WIND'; val = 'WIND'; }, - { key = 'D_RH'; val = 'RH'; }, - { key = 'D_MIXR'; val = 'MIXR'; }, - { key = 'D_PRMSL'; val = 'PRMSL'; }, - { key = 'D_PBL'; val = 'PBL'; }, - { key = 'D_CAPE'; val = 'CAPE'; } - { key = 'D_MLCAPE'; val = 'MLCAPE'; } - ]; - -The BUFR variable names are not shared with other forecast data. This map is used to convert the BUFR name to the common name, like GRIB2. It allows to share the configuration for forecast data with PB2NC observation data. If there is no mapping, the BUFR variable name will be saved to output NetCDF file. + obs_bufr_map = [ + { key = 'POB'; val = 'PRES'; }, + { key = 'QOB'; val = 'SPFH'; }, + { key = 'TOB'; val = 'TMP'; }, + { key = 'ZOB'; val = 'HGT'; }, + { key = 'UOB'; val = 'UGRD'; }, + { key = 'VOB'; val = 'VGRD'; }, + { key = 'D_DPT'; val = 'DPT'; }, + { key = 'D_WDIR'; val = 'WDIR'; }, + { key = 'D_WIND'; val = 'WIND'; }, + { key = 'D_RH'; val = 'RH'; }, + { key = 'D_MIXR'; val = 'MIXR'; }, + { key = 'D_PRMSL'; val = 'PRMSL'; }, + { key = 'D_PBL'; val = 'PBL'; }, + { key = 'D_CAPE'; val = 'CAPE'; }, + { key = 'D_MLCAPE'; val = 'MLCAPE'; } + ]; + +The BUFR variable names are not shared with other forecast data. This map is used to convert the BUFR name to the common name, like GRIB2. It allows the configuration for forecast data to be shared with PB2NC observation data. If there is no mapping, the BUFR variable name will be saved to the output NetCDF file. _____________________ .. code-block:: none - quality_mark_thresh = <=2; + quality_mark_thresh = <=2; -Each observation has an integer quality mark value associated with it. The **quality_mark_thresh** is used to stratify which quality marks will be retained. By default, observations with quality marks less than or equal to 2 will be kept. This can be specified as a threshold string (e.g. :code:`<=2||==9` for less than or equal to 2 or exactly equal to 9) or as an integer defining the maximum allowable quality mark value. Earlier versions of MET only supported the integer setting. +Each observation has an integer quality mark value associated with it. The **quality_mark_thresh** is used to stratify which quality marks will be retained. By default, observations with quality marks less than or equal to 2 will be kept. This can be specified as a threshold string (e.g., :code:`<=2||==9` for less than or equal to 2 or exactly equal to 9) or as an integer defining the maximum allowable quality mark value. Earlier versions of MET only supported the integer setting. _____________________ .. code-block:: none - event_stack_flag = TOP; + event_stack_flag = TOP; -A PrepBUFR message may contain duplicate observations with different quality mark values. The **event_stack_flag** indicates whether to use the observations at the top of the event stack (observation values have had more quality control processing applied) or the bottom of the event stack (observation values have had no quality control processing applied). The flag value of **TOP** listed above indicates the observations with the most amount of quality control processing should be used, the **BOTTOM** option uses the data closest to raw values. +A PrepBUFR message may contain duplicate observations with different quality mark values. The **event_stack_flag** indicates whether to use the observations at the top of the event stack (observation values have had more quality control processing applied) or the bottom of the event stack (observation values have had no quality control processing applied). The flag value of **TOP** listed above indicates the observations with the most amount of quality control processing should be used; the **BOTTOM** option uses the data closest to raw values. _____________________ .. code-block:: none - time_summary = { - flag = FALSE; - raw_data = FALSE; - beg = "000000"; - end = "235959"; - step = 300; - width = 600; - // width = { beg = -300; end = 300; } - grib_code = []; - obs_var = [ "TMP", "WDIR", "RH" ]; - type = [ "min", "max", "range", "mean", "stdev", "median", "p80" ]; - vld_freq = 0; - vld_thresh = 0.0; - } + time_summary = { + flag = FALSE; + raw_data = FALSE; + beg = "000000"; + end = "235959"; + step = 300; + width = 600; + // width = { beg = -300; end = 300; } + grib_code = []; + obs_var = [ "TMP", "WDIR", "RH" ]; + type = [ "min", "max", "range", "mean", "stdev", "median", "p80" ]; + vld_freq = 0; + vld_thresh = 0.0; + } The **time_summary** dictionary enables additional processing for observations with high temporal resolution. The **flag** entry toggles the **time_summary** on (**TRUE**) and off (**FALSE**). If the **raw_data** flag is set to TRUE, then both the individual observation values and the derived time summary value will be written to the output. If FALSE, only the summary values are written. Observations may be summarized across the user specified time period defined by the **beg** and **end** entries in HHMMSS format. The **step** entry defines the time between intervals in seconds. The **width** entry specifies the summary interval in seconds. It may either be set as an integer number of seconds for a centered time interval or a dictionary with beginning and ending time offsets in seconds. -This example listed above does a 10-minute time summary (width = 600;) every 5 minutes (step = 300;) throughout the day (beg = "000000"; end = 235959";). The first interval will be from 23:55:00 the previous day through 00:04:59 of the current day. The second interval will be from 0:00:00 through 00:09:59. And so on. +This example listed above does a 10-minute time summary (width = 600;) every 5 minutes (step = 300;) throughout the day (beg = "000000"; end = "235959";). The first interval will be from 23:55:00 the previous day through 00:04:59 of the current day. The second interval will be from 0:00:00 through 00:09:59. And so on. The two **width** settings listed above are equivalent. Both define a centered 10-minute time interval. Use the **beg** and **end** entries to define uncentered time intervals. The following example requests observations for one hour prior: .. code-block:: none - width = { beg = -3600; end = 0; } + width = { beg = -3600; end = 0; } -The summaries will only be calculated for the observations specified in the **grib_code** or **obs_var** entries. The **grib_code** entry is an array of integers while the **obs_var** entries is an array of strings. The supported summaries are **min** (minimum), **max** (maximum), **range, mean, stdev** (standard deviation), **median** and **p##** (percentile, with the desired percentile value specified in place of ##). If multiple summaries are selected in a single run, a string indicating the summary method applied will be appended to the output message type. +The summaries will only be calculated for the observations specified in the **grib_code** or **obs_var** entries. The **grib_code** entry is an array of integers while the **obs_var** entry is an array of strings. The supported summaries are **min** (minimum), **max** (maximum), **range, mean, stdev** (standard deviation), **median** and **p##** (percentile, with the desired percentile value specified in place of ##). If multiple summaries are selected in a single run, a string indicating the summary method applied will be appended to the output message type. The **vld_freq** and **vld_thresh** entries specify the required ratio of valid data for an output time summary value to be computed. This option is only applied when these entries are set to non-zero values. The **vld_freq** entry specifies the expected frequency of observations in seconds. The width of the time window is divided by this frequency to compute the expected number of observations for the time window. The actual number of valid observations is divided by the expected number to compute the ratio of valid data. An output time summary value will only be written if that ratio is greater than or equal to the **vld_thresh** entry. Detailed information about which observations are excluded is provided at debug level 4. @@ -311,127 +311,127 @@ Each NetCDF file generated by the PB2NC tool contains the dimensions and variabl .. _table_reformat-point_pb2nc_output_dim: .. list-table:: NetCDF file dimensions for pb2nc output - :widths: auto - :header-rows: 1 - - * - NetCDF Dimension - - Description - * - mxstr, mxstr2, mxstr3 - - Maximum string lengths (16, 40, and 80) - * - nobs - - Number of PrepBUFR observations in the file (UNLIMITED) - * - nhdr, npbhdr - - Number of PrepBUFR messages in the file (variable) - * - nhdr_typ, nhdr_sid, nhdr_vld - - Number of unique header message type, station ID, and valid time strings (variable) - * - nobs_qty - - Number of unique quality control strings (variable) - * - obs_var_num - - Number of unique observation variable types (variable) + :widths: auto + :header-rows: 1 + + * - NetCDF Dimension + - Description + * - mxstr, mxstr2, mxstr3 + - Maximum string lengths (16, 40, and 80) + * - nobs + - Number of PrepBUFR observations in the file (UNLIMITED) + * - nhdr, npbhdr + - Number of PrepBUFR messages in the file (variable) + * - nhdr_typ, nhdr_sid, nhdr_vld + - Number of unique header message type, station ID, and valid time strings (variable) + * - nobs_qty + - Number of unique quality control strings (variable) + * - obs_var_num + - Number of unique observation variable types (variable) .. _table_reformat-point_pb2nc_output_vars: .. list-table:: NetCDF variables in pb2nc output - :widths: auto - :header-rows: 1 - - * - NetCDF Variable - - Dimension - - Description - - Data Type - * - obs_qty - - nobs - - Integer value of the n_obs_qty dimension for the observation quality control string - - Integer - * - obs_hid - - nobs - - Integer value of the nhdr dimension for the header arrays with which this observation is associated - - Integer - * - obs_vid - - nobs - - Integer value of the obs_var_num dimension for the observation variable name, units, and description - - Integer - * - obs_lvl - - nobs - - Floating point pressure level in hPa or accumulation interval - - Float - * - obs_hgt - - nobs - - Floating point height in meters above sea level - - Float - * - obs_val - - nobs - - Floating point observation value. - - Float - * - hdr_typ - - nhdr - - Integer value of the nhdr_typ dimension for the message type string - - Integer - * - hdr_sid - - nhdr - - Integer value of the nhdr_sid dimension for the station ID string - - Integer - * - hdr_vld - - nhdr - - Integer value of the nhdr_vld dimension for the valid time string - - Integer - * - hdr_lat, hdr_lon - - nhdr - - Floating point latitude in degrees north and longitude in degrees east - - Float - * - hdr_elv - - nhdr - - Floating point elevation of observing station in meters above sea level - - Float - * - hdr_prpt_typ - - npbhdr - - Integer PrepBUFR report type value - - Integer - * - hdr_irpt_typ - - npbhdr - - Integer input report type value - - Integer - * - hdr_inst_typ - - npbhdr - - Integer instrument type value - - Integer - * - hdr_typ_table - - nhdr_typ, - - mxstr2 Lookup table containing unique message type strings - - String - * - hdr_sid_table - - nhdr_sid, mxstr2 - - mxstr2 Lookup table containing unique station ID strings - - String - * - hdr_vld_table - - nhdr_vld, mxstr - - Lookup table containing unique valid time strings in YYYYMMDD_HHMMSS UTC format - - Datetime String - * - obs_qty_table - - nobs_qty, mxstr - - Lookup table containing unique quality control strings - - String - * - obs_var - - obs_var_num, mxstr - - Lookup table containing unique observation variable names - - String - * - obs_unit - - obs_var_num, mxstr2 - - Lookup table containing a units string for the unique observation variable names in obs_var - - String - * - obs_desc - - obs_var_num, mxstr3 - - Lookup table containing a description string for the unique observation variable names in obs_var - - String + :widths: auto + :header-rows: 1 + + * - NetCDF Variable + - Dimension + - Description + - Data Type + * - obs_qty + - nobs + - Integer value of the n_obs_qty dimension for the observation quality control string + - Integer + * - obs_hid + - nobs + - Integer value of the nhdr dimension for the header arrays with which this observation is associated + - Integer + * - obs_vid + - nobs + - Integer value of the obs_var_num dimension for the observation variable name, units, and description + - Integer + * - obs_lvl + - nobs + - Floating point pressure level in hPa or accumulation interval + - Float + * - obs_hgt + - nobs + - Floating point height in meters above sea level + - Float + * - obs_val + - nobs + - Floating point observation value. + - Float + * - hdr_typ + - nhdr + - Integer value of the nhdr_typ dimension for the message type string + - Integer + * - hdr_sid + - nhdr + - Integer value of the nhdr_sid dimension for the station ID string + - Integer + * - hdr_vld + - nhdr + - Integer value of the nhdr_vld dimension for the valid time string + - Integer + * - hdr_lat, hdr_lon + - nhdr + - Floating point latitude in degrees north and longitude in degrees east + - Float + * - hdr_elv + - nhdr + - Floating point elevation of observing station in meters above sea level + - Float + * - hdr_prpt_typ + - npbhdr + - Integer PrepBUFR report type value + - Integer + * - hdr_irpt_typ + - npbhdr + - Integer input report type value + - Integer + * - hdr_inst_typ + - npbhdr + - Integer instrument type value + - Integer + * - hdr_typ_table + - nhdr_typ, mxstr2 + - Lookup table containing unique message type strings + - String + * - hdr_sid_table + - nhdr_sid, mxstr2 + - Lookup table containing unique station ID strings + - String + * - hdr_vld_table + - nhdr_vld, mxstr + - Lookup table containing unique valid time strings in YYYYMMDD_HHMMSS UTC format + - Datetime String + * - obs_qty_table + - nobs_qty, mxstr + - Lookup table containing unique quality control strings + - String + * - obs_var + - obs_var_num, mxstr + - Lookup table containing unique observation variable names + - String + * - obs_unit + - obs_var_num, mxstr2 + - Lookup table containing a units string for the unique observation variable names in obs_var + - String + * - obs_desc + - obs_var_num, mxstr3 + - Lookup table containing a description string for the unique observation variable names in obs_var + - String ASCII2NC Tool ============= This section describes how to run the ASCII2NC tool. The ASCII2NC tool is used to reformat ASCII point observations into the NetCDF format expected by the Point-Stat tool. For those users wishing to verify against point observations that are not available in PrepBUFR format, the ASCII2NC tool provides a way of incorporating those observations into MET. If the ASCII2NC tool is used to perform a reformatting step, no configuration file is needed. However, for more complex processing, such as summarizing time series observations, a configuration file may be specified. For details on the configuration file options, see :numref:`config_options` and example configuration files distributed with the MET code. -While initial versions of the ASCII2NC tool only supported a simple 11 column ASCII point observation format, support for several additional formats has been added. It currently supports point observation data in the following formats: +While initial versions of the ASCII2NC tool only supported a simple 11-column ASCII point observation format, support for several additional formats has been added. It currently supports point observation data in the following formats: -• Default 11 column MET point observation format, as described in :numref:`table_reformat-point_ascii2nc_format` +• Default 11-column MET point observation format, as described in :numref:`table_reformat-point_ascii2nc_format` • `little_r format `_ @@ -469,7 +469,7 @@ The default ASCII point observation format consists of one row of data per obser - Text string containing the observation message type as described in the previous section on the PB2NC tool (max 40 characters). * - 2 - Station_ID - - Text string containing the station id (max 40 characters). + - Text string containing the station ID (max 40 characters). * - 3 - Valid_Time - Text string containing the observation valid time in YYYYMMDD_HHMMSS format. @@ -501,7 +501,7 @@ The default ASCII point observation format consists of one row of data per obser ascii2nc Usage -------------- -Once the ASCII point observations have been formatted as expected, the ASCII file is ready to be processed by the ASCII2NC tool. The usage statement for ASCII2NC tool is shown below: +Once the ASCII point observations have been formatted as expected, the ASCII file is ready to be processed by the ASCII2NC tool. The usage statement for the ASCII2NC tool is shown below: .. code-block:: none @@ -529,7 +529,7 @@ Required Arguments for ascii2nc - a regular file - an ASCII file list, as described in :numref:`ascii_file_lists` - - a top-level diretory to be recursively searched for files matching the **-inputrx reg_exp**, described below + - a top-level directory to be recursively searched for files matching the **-inputrx reg_exp**, described below If using Python embedding with the **-format python** option, specify inputs as quoted strings containing the Python script to be run followed by any command line arguments for that script. @@ -552,7 +552,7 @@ Optional Arguments for ascii2nc 9. The **-mask_poly** file option is a polyline masking file to filter the point observations spatially. -10. The **-mask_sid** file|list option is a station ID masking file or a comma-separated list of station ID's to filter the point observations spatially. See the description of the "sid" entry in :numref:`config_options`. +10. The **-mask_sid** file|list option is a station ID masking file or a comma-separated list of station IDs to filter the point observations spatially. See the description of the "sid" entry in :numref:`config_options`. 11. The **-log file** option directs output and errors to the specified log file. All messages will be written to that file as well as standard out and error. Thus, users can save the messages without having to redirect the output on the command line. The default behavior is no log file. @@ -564,8 +564,8 @@ An example of the ascii2nc calling sequence is shown below: .. code-block:: none - ascii2nc sample_ascii_obs.txt \ - sample_ascii_obs.nc + ascii2nc sample_ascii_obs.txt \ + sample_ascii_obs.nc In this example, the ASCII2NC tool will reformat the input **sample_ascii_obs.txt file** into NetCDF format and write the output to a file named **sample_ascii_obs.nc**. @@ -582,7 +582,7 @@ _____________________ .. code-block:: none - version = "VN.N"; + version = "VN.N"; The configuration options listed above are common to many MET tools and are described in :numref:`config_options`. @@ -590,7 +590,7 @@ _____________________ .. code-block:: none - time_summary = { ... } + time_summary = { ... } The **time_summary** feature was implemented to allow additional processing of observations with high temporal resolution, such as SURFRAD data every 5 minutes. This option is described in :numref:`pb2nc configuration file`. @@ -598,17 +598,17 @@ _____________________ .. code-block:: none - message_type_map = [ - { key = "FM-12 SYNOP"; val = "ADPSFC"; }, - { key = "FM-13 SHIP"; val = "SFCSHP"; }, - { key = "FM-15 METAR"; val = "ADPSFC"; }, - { key = "FM-18 BUOY"; val = "SFCSHP"; }, - { key = "FM-281 QSCAT"; val = "ASCATW"; }, - { key = "FM-32 PILOT"; val = "ADPUPA"; }, - { key = "FM-35 TEMP"; val = "ADPUPA"; }, - { key = "FM-88 SATOB"; val = "SATWND"; }, - { key = "FM-97 ACARS"; val = "AIRCFT"; } - ]; + message_type_map = [ + { key = "FM-12 SYNOP"; val = "ADPSFC"; }, + { key = "FM-13 SHIP"; val = "SFCSHP"; }, + { key = "FM-15 METAR"; val = "ADPSFC"; }, + { key = "FM-18 BUOY"; val = "SFCSHP"; }, + { key = "FM-281 QSCAT"; val = "ASCATW"; }, + { key = "FM-32 PILOT"; val = "ADPUPA"; }, + { key = "FM-35 TEMP"; val = "ADPUPA"; }, + { key = "FM-88 SATOB"; val = "SATWND"; }, + { key = "FM-97 ACARS"; val = "AIRCFT"; } + ]; This entry is an array of dictionaries, each containing a **key** string and **val** string which define a mapping of input strings to output message types. This mapping is currently only applied when converting input little_r report types to output message types. @@ -617,7 +617,7 @@ ascii2nc Output The NetCDF output of the ASCII2NC tool is structured in the same way as the output of the PB2NC tool described in :numref:`pb2nc output`. -"obs_vid" variable is replaced with "obs_gc" when the GRIB code is given instead of the variable names. In this case, the global attribute "use_var_id" does not exist or set to false (use_var_id = "false" ;). Three variables (obs_var, obs_units, and obs_desc) related with variable names are not added. +"obs_vid" variable is replaced with "obs_gc" when the GRIB code is given instead of the variable names. In this case, the global attribute "use_var_id" does not exist or is set to false (use_var_id = "false" ;). Three variables (obs_var, obs_units, and obs_desc) related to variable names are not added. MADIS2NC Tool ============= @@ -663,7 +663,7 @@ Optional Arguments for madis2nc 4. The **-config file** option specifies the configuration file to generate summaries of the fields in the ASCII files. -5. The **-qc_dd list** option specifies a comma-separated list of QC flag values to be accepted(Z,C,S,V,X,Q,K,G,B). +5. The **-qc_dd list** option specifies a comma-separated list of QC flag values to be accepted (Z,C,S,V,X,Q,K,G,B). 6. The **-lvl_dim list** option specifies a comma-separated list of vertical level dimensions to be processed. @@ -673,7 +673,7 @@ Optional Arguments for madis2nc 9. The **-mask_poly file** option defines a polyline masking file for filtering the point observations spatially. -10. The **-mask_sid file|list** option is a station ID masking file or a comma-separated list of station ID's for filtering the point observations spatially. See the description of the "sid" entry in :numref:`config_options`. +10. The **-mask_sid file|list** option is a station ID masking file or a comma-separated list of station IDs for filtering the point observations spatially. See the description of the "sid" entry in :numref:`config_options`. 11. The **-log file** option directs output and errors to the specified log file. All messages will be written to that file as well as standard out and error. Thus, users can save the messages without having to redirect the output on the command line. The default behavior is no log file. @@ -685,8 +685,8 @@ An example of the madis2nc calling sequence is shown below: .. code-block:: none - madis2nc sample_madis_obs.nc \ - sample_madis_obs_met.nc -log madis.log -v 3 + madis2nc sample_madis_obs.nc \ + sample_madis_obs_met.nc -log madis.log -v 3 In this example, the MADIS2NC tool will reformat the input sample_madis_obs.nc file into NetCDF format and write the output to a file named sample_madis_obs_met.nc. Warnings and error messages will be written to the madis.log file, and the verbosity level of logging is three. @@ -701,7 +701,7 @@ _____________________ .. code-block:: none - version = "VN.N"; + version = "VN.N"; The configuration options listed above are common to many MET tools and are described in :numref:`config_options`. @@ -709,7 +709,7 @@ _____________________ .. code-block:: none - time_summary = { ... } + time_summary = { ... } The **time_summary** dictionary is described in :numref:`pb2nc configuration file`. @@ -717,29 +717,29 @@ _____________________ .. code-block:: none - grib_var_map = [ - { key = "1" ; val = "PRES,Pa" ; }, // Station Pressure - { key = "2" ; val = "PRMSL,Pa" ; }, // Sea Level Pressure - { key = "7" ; val = "HGT,gpm" ; }, // Height - { key = "11" ; val = "TMP,K" ; }, // Temperature - { key = "15" ; val = "TMAX,K" ; }, // Maximum Temperature - { key = "16" ; val = "TMIN,K" ; }, // Minimum Temperature - { key = "17" ; val = "DPT,K" ; }, // Dewpoint - { key = "20" ; val = "VISIB,W/m^2" ; }, // Visibility - { key = "31" ; val = "WDIR,deg" ; }, // Wind Direction - { key = "32" ; val = "WIND,m/s" ; }, // Wind Speed - { key = "33" ; val = "UGRD,m/s" ; }, // Write U-component of wind - { key = "34" ; val = "VGRD,m/s" ; }, // Write V-component of wind - { key = "52" ; val = "RH,%" ; }, // Relative Humidity - { key = "54" ; val = "PWAT,kg/m^2" ; }, // Precipitable Water - { key = "59" ; val = "PRATE,kg/m^2/s"; }, // Precipitation Rate - { key = "61" ; val = "APCP,kg/m^2" ; }, // Precipitation - { key = "66" ; val = "SNOD,m" ; }, // Snow Cover - { key = "80" ; val = "WTMP,K" ; }, // Sea Surface Temperature - { key = "85" ; val = "TSOIL,K" ; }, // Soil Temperature - { key = "180"; val = "GUST,m/s" ; }, // Wind Gust - { key = "250"; val = "SWHR,K/s" ; } // Solar Radiation - ]; + grib_var_map = [ + { key = "1" ; val = "PRES,Pa" ; }, // Station Pressure + { key = "2" ; val = "PRMSL,Pa" ; }, // Sea Level Pressure + { key = "7" ; val = "HGT,gpm" ; }, // Height + { key = "11" ; val = "TMP,K" ; }, // Temperature + { key = "15" ; val = "TMAX,K" ; }, // Maximum Temperature + { key = "16" ; val = "TMIN,K" ; }, // Minimum Temperature + { key = "17" ; val = "DPT,K" ; }, // Dewpoint + { key = "20" ; val = "VISIB,W/m^2" ; }, // Visibility + { key = "31" ; val = "WDIR,deg" ; }, // Wind Direction + { key = "32" ; val = "WIND,m/s" ; }, // Wind Speed + { key = "33" ; val = "UGRD,m/s" ; }, // Write U-component of wind + { key = "34" ; val = "VGRD,m/s" ; }, // Write V-component of wind + { key = "52" ; val = "RH,%" ; }, // Relative Humidity + { key = "54" ; val = "PWAT,kg/m^2" ; }, // Precipitable Water + { key = "59" ; val = "PRATE,kg/m^2/s"; }, // Precipitation Rate + { key = "61" ; val = "APCP,kg/m^2" ; }, // Precipitation + { key = "66" ; val = "SNOD,m" ; }, // Snow Cover + { key = "80" ; val = "WTMP,K" ; }, // Sea Surface Temperature + { key = "85" ; val = "TSOIL,K" ; }, // Soil Temperature + { key = "180"; val = "GUST,m/s" ; }, // Wind Gust + { key = "250"; val = "SWHR,K/s" ; } // Solar Radiation + ]; The GRIB code mappings for variable names and units are defined in the **grib_var_map** dictionary within **Madis2NcConfig_default**. In this mapping, each key is a GRIB code, and the corresponding value is a pair: the first element is the variable name, and the second is the unit. @@ -748,7 +748,7 @@ madis2nc Output The NetCDF output of the MADIS2NC tool is structured in the same way as the output of the PB2NC tool described in :numref:`pb2nc output`. -"obs_vid" variable is replaced with "obs_gc" when the GRIB code is given instead of the variable names. In this case, the global attribute "use_var_id" does not exist or set to false (use_var_id = "false" ;). Three variables (obs_var, obs_units, and obs_desc) related with variable names are not added. +"obs_vid" variable is replaced with "obs_gc" when the GRIB code is given instead of the variable names. In this case, the global attribute "use_var_id" does not exist or is set to false (use_var_id = "false" ;). Three variables (obs_var, obs_units, and obs_desc) related to variable names are not added. Starting from MET version 12.2.0, the GRIB codes were replaced with variable names which are defined in the **grib_var_map** dictionary within **Madis2NcConfig_default**. @@ -761,7 +761,7 @@ The LIDAR2NC tool creates a NetCDF point observation file from a CALIPSO HDF dat lidar2nc Usage -------------- -The usage statement for LIDAR2NC tool is shown below: +The usage statement for the LIDAR2NC tool is shown below: .. code-block:: none @@ -898,13 +898,13 @@ Required Arguments for ioda2nc Optional Arguments for ioda2nc ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ -3. The **-config config_file** is a IODA2NCConfig file to filter the point observations and define time summaries. +3. The **-config config_file** is an IODA2NCConfig file to filter the point observations and define time summaries. -4. The **-obs_var var_list** setting is a comma-separated list of variables to be saved from input the input file (by defaults, saves "all"). +4. The **-obs_var var_list** setting is a comma-separated list of variables to be saved from the input file (by default, saves "all"). 5. The **-iodafile ioda_file** option specifies additional input IODA observation files to be processed. -6. The **-valid_beg time** and **-valid_end time** options in YYYYMMDD[_HH[MMSS]] format overrides the retention time window from the configuration file. +6. The **-valid_beg time** and **-valid_end time** options in YYYYMMDD[_HH[MMSS]] format override the retention time window from the configuration file. 7. The **-nmsg n** indicates the number of IODA records to process. @@ -918,27 +918,27 @@ An example of the ioda2nc calling sequence is shown below: .. code-block:: none - ioda2nc \ - ioda.NC001007.2020031012.nc ioda2nc.2020031012.nc \ - -config IODA2NCConfig -v 3 -lg run_ioda2nc.log + ioda2nc \ + ioda.NC001007.2020031012.nc ioda2nc.2020031012.nc \ + -config IODA2NCConfig -v 3 -log run_ioda2nc.log -In this example, the IODA2NC tool will reformat the data in the input ioda.NC001007.2020031012.nc file and write the output to a file named ioda2nc.2020031012.nc. The data to be processed is specified by IODA2NCConfig, log messages will be written to the ioda2nc.log file, and the verbosity level is three. +In this example, the IODA2NC tool will reformat the data in the input ioda.NC001007.2020031012.nc file and write the output to a file named ioda2nc.2020031012.nc. The data to be processed is specified by IODA2NCConfig, log messages will be written to the run_ioda2nc.log file, and the verbosity level is three. ioda2nc Configuration File -------------------------- The default configuration file for the IODA2NC tool named **IODA2NcConfig_default** can be found in the installed *share/met/config* directory. It is recommended that users make a copy of this file prior to modifying its contents. -The IODA2NC configuration file is optional and only necessary when defining filtering the input observations or defining time summaries. The contents of the default IODA2NC configuration file are described below. +The IODA2NC configuration file is optional and only necessary when filtering the input observations or defining time summaries. The contents of the default IODA2NC configuration file are described below. _____________________ .. code-block:: none - obs_window = { beg = -5400; end = 5400; } - mask = { grid = ""; poly = ""; } - tmp_dir = "/tmp"; - version = "VN.N"; + obs_window = { beg = -5400; end = 5400; } + mask = { grid = ""; poly = ""; } + tmp_dir = "/tmp"; + version = "VN.N"; The configuration options listed above are common to many MET tools and are described in :numref:`config_options`. @@ -946,28 +946,28 @@ _____________________ .. code-block:: none - message_type = []; - message_type_group_map = []; - message_type_map = []; - station_id = []; - elevation_range = { ... }; - level_range = { ... }; - obs_var = []; - quality_mark_thresh = NA; - time_summary = { ... } + message_type = []; + message_type_group_map = []; + message_type_map = []; + station_id = []; + elevation_range = { ... }; + level_range = { ... }; + obs_var = []; + quality_mark_thresh = NA; + time_summary = { ... } The configuration options listed above are supported by other point observation pre-processing tools and are described in :numref:`pb2nc configuration file`. .. note:: - Setting "quality_mark_thresh" as an NA threshold, as shown above, always evaluates to true. - So by default, IODA2NC performs no filtering of observations based on their quality mark value. + Setting "quality_mark_thresh" as an NA threshold, as shown above, always evaluates to true. + So by default, IODA2NC performs no filtering of observations based on their quality mark value. _____________________ .. code-block:: none - obs_name_map = []; + obs_name_map = []; This entry is an array of dictionaries, each containing a **key** string and **val** string which define a mapping of input IODA variable names to output variable names. The default IODA map, obs_var_map, is appended to this map. @@ -975,26 +975,26 @@ _____________________ .. code-block:: none - metadata_map = [ - { key = "message_type"; val = "msg_type,station_ob"; }, - { key = "station_id"; val = "station_id,report_identifier"; }, - { key = "pressure"; val = "air_pressure,pressure"; }, - { key = "height"; val = "height,height_above_mean_sea_level"; }, - { key = "elevation"; val = "elevation,station_elevation"; }, - { key = "nlocs"; val = "Location"; } - ]; + metadata_map = [ + { key = "message_type"; val = "msg_type,station_ob"; }, + { key = "station_id"; val = "station_id,report_identifier"; }, + { key = "pressure"; val = "air_pressure,pressure"; }, + { key = "height"; val = "height,height_above_mean_sea_level"; }, + { key = "elevation"; val = "elevation,station_elevation"; }, + { key = "nlocs"; val = "Location"; } + ]; This entry is an array of dictionaries, each containing a **key** string and **val** string which define a mapping of metadata for IODA data files. -The "nlocs" is for the dimension name of the locations. The following key can be added: "nstring", "latitude" and "longitude". +The "nlocs" is for the dimension name of the locations. The following keys can be added: "nstring", "latitude" and "longitude". _____________________ .. code-block:: none - obs_to_qc_map = [ - { key = "wind_from_direction"; val = "eastward_wind,northward_wind"; }, - { key = "wind_speed"; val = "eastward_wind,northward_wind"; } - ]; + obs_to_qc_map = [ + { key = "wind_from_direction"; val = "eastward_wind,northward_wind"; }, + { key = "wind_speed"; val = "eastward_wind,northward_wind"; } + ]; This entry is an array of dictionaries, each containing a **key** string and **val** string which define a mapping of QC variable name for IODA data files. @@ -1002,7 +1002,7 @@ _____________________ .. code-block:: none - missing_thresh = [ <=-1e9, >=1e9, ==-9999 ]; + missing_thresh = [ <=-1e9, >=1e9, ==-9999 ]; The **missing_thresh** option is an array of thresholds. Any data values which meet any of these thresholds are interpreted as being bad, or missing, data. @@ -1014,7 +1014,7 @@ The NetCDF output of the IODA2NC tool is structured in the same way as the outpu Point2Grid Tool =============== -The Point2Grid tool reads point observations from a MET NetCDF point obseravtion file, via python embedding, or from GOES NetCDF input files (especially, Aerosol Optical Depth) and creates a gridded NetCDF file. Future development may add support for additional input types. +The Point2Grid tool reads point observations from a MET NetCDF point observation file, via Python embedding, or from GOES NetCDF input files (especially, Aerosol Optical Depth) and creates a gridded NetCDF file. Future development may add support for additional input types. point2grid Usage ---------------- @@ -1047,13 +1047,13 @@ Required Arguments for point2grid 1. The **input_filename** argument indicates the name of the input file to be processed. The input can be a MET NetCDF point observation file generated by other MET tools or a GOES NetCDF AOD dataset. Python embedding for point observations is also supported, as described in :numref:`pyembed-point-obs-data`. -The MET point observation NetCDF file name as **input_filename** argument is equivalent with "PYTHON_NUMPY=MET_BASE/python/examples/read_met_point_obs.py netcdf_filename". +The MET point observation NetCDF file name as **input_filename** argument is equivalent to "PYTHON_NUMPY=MET_BASE/python/examples/read_met_point_obs.py netcdf_filename". 2. The **to_grid** argument defines the output grid as: (1) a named grid, (2) the path to a gridded data file, or (3) an explicit grid specification string. 3. The **output_filename** argument is the name of the output NetCDF file to be written. -4. The **-field** string argument is a string that defines the data to be regridded. It may be used multiple times. If **-adp** option is given (for GOES AOD data), the name consists with the variable name from the input data file and the variable name from ADP data file (for example, "AOD_Smoke" or "AOD_Dust": getting AOD variable from the input data and applying smoke or dust variable from ADP data file). +4. The **-field** string argument is a string that defines the data to be regridded. It may be used multiple times. If **-adp** option is given (for GOES AOD data), the name consists of the variable name from the input data file and the variable name from ADP data file (for example, "AOD_Smoke" or "AOD_Dust": getting AOD variable from the input data and applying smoke or dust variable from ADP data file). Optional Arguments for point2grid ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ @@ -1064,7 +1064,7 @@ Optional Arguments for point2grid 7. The **-goes_qc** flags option specifies a comma-separated list of quality control (QC) flags, for example "0,1". Only used if grid_mapping is set to "goes_imager_projection" and the QC variable exists. Note that the older **-qc** option name is also supported. -9. The **-adp adp_filename** option provides an additional Aerosol Detection Product (ADP) information on aerosols, dust, and smoke. This option is ignored if the requested variable is not GOES AOD ("AOD_Dust" or "AOD_Smoke"). The gridded data is filtered by the presence of dust/smoke. If -goes_qc options are given, it's applied to QC of dust/smoke, too (First filtering with AOD QC values and the second filtering with dust/smoke QC values). +8. The **-adp adp_filename** option provides an additional Aerosol Detection Product (ADP) information on aerosols, dust, and smoke. This option is ignored if the requested variable is not GOES AOD ("AOD_Dust" or "AOD_Smoke"). The gridded data is filtered by the presence of dust/smoke. If -goes_qc options are given, it's applied to QC of dust/smoke, too (First filtering with AOD QC values and the second filtering with dust/smoke QC values). 9. The **-method type** option specifies the regridding method. The default method is UW_MEAN. @@ -1090,39 +1090,39 @@ For the GOES-East and GOES-West data, computing the latitude and longitude pixel .. code-block:: none - point2grid \ - OR_ABI-L2-AODC-M3_G16_s20181341702215_e20181341704588_c20181341711418.nc \ - G212 \ - regrid_data_plane_GOES-16_AOD_TO_G212.nc \ - -field 'name="AOD"; level="(*,*)";' \ - -goes_qc 0,1,2 \ - -method MAX + point2grid \ + OR_ABI-L2-AODC-M3_G16_s20181341702215_e20181341704588_c20181341711418.nc \ + G212 \ + regrid_data_plane_GOES-16_AOD_TO_G212.nc \ + -field 'name="AOD"; level="(*,*)";' \ + -goes_qc 0,1,2 \ + -method MAX -When processing GOES data, the **-goes_qc** option may also be used to specify the acceptable quality control flag values. The example above regrids the GOES-East AOD values to NCEP Grid number 212 (which QC flags are high, medium, and low), writing to the output the maximum AOD value falling inside each grid box. +When processing GOES data, the **-goes_qc** option may also be used to specify the acceptable quality control flag values. The example above regrids the GOES-East AOD values to NCEP Grid number 212 using only pixels whose QC flags are 0, 1, or 2 (high, medium, and low quality), writing to the output the maximum AOD value falling inside each grid box. The grid name or the grid definition can be given with the -field option when the grid information is missing from the input NetCDF file for the latitude_longitude projection. The latitude and longitude variable names should be defined by the user, and the grid information from the set_attr_grid is ignored in this case except nx and ny. .. code-block:: none - point2grid \ - iceh.2018-01-03.c00.tlat_tlon.nc \ - G231 \ - point2grid_cice_to_G231.nc \ - -config Point2GridConfig_tlat_tlon \ - -field 'name="hi_d"; level="(0,*,*)"; set_attr_grid="latlon 1440 1080 -79.80672 60.28144 0.04 0.04";' + point2grid \ + iceh.2018-01-03.c00.tlat_tlon.nc \ + G231 \ + point2grid_cice_to_G231.nc \ + -config Point2GridConfig_tlat_tlon \ + -field 'name="hi_d"; level="(0,*,*)"; set_attr_grid="latlon 1440 1080 -79.80672 60.28144 0.04 0.04";' Listed below is an example of using Python embedding to pass point observations as input to point2grid: .. code-block:: none - point2grid \ - 'PYTHON_NUMPY=MET_BASE/python/examples/read_met_point_obs.py ascii2nc_edr_hourly.20130827.nc' \ - G212 \ - python_gridded_ascii_python.nc -config Point2GridConfig_edr \ - -field 'name="200"; level="*"; valid_time="20130827_205959";' \ - -method MAX + point2grid \ + 'PYTHON_NUMPY=MET_BASE/python/examples/read_met_point_obs.py ascii2nc_edr_hourly.20130827.nc' \ + G212 \ + python_gridded_ascii_python.nc -config Point2GridConfig_edr \ + -field 'name="200"; level="*"; valid_time="20130827_205959";' \ + -method MAX Please refer to :numref:`Appendix F, Section %s ` for more details about Python embedding in MET. @@ -1138,7 +1138,7 @@ The point2grid tool will output a gridded NetCDF file containing the following: 3. The variable specified in the -field string regridded to the grid defined in the **to_grid** argument. -4. The count field which represents the number of point observations that were included calculating the value of the variable at that grid cell. +4. The count field which represents the number of point observations that were included in calculating the value of the variable at that grid cell. 5. The mask field which is a binary field representing the presence or lack thereof of point observations at that grid cell. A value of "1" indicates that there was at least one point observation within the bounds of that grid cell and a value of "0" indicates the lack of point observations at that grid cell. @@ -1146,7 +1146,7 @@ The point2grid tool will output a gridded NetCDF file containing the following: 7. The probability mask field which is a binary field that represents whether or not there is probability data at that grid point. Can be either "0" or "1" with "0" meaning the probability value does not exist and a value of "1" meaning that the probability value does exist. -For MET observation input and CF complaint NetCDF input with 2D time variable: The latest observation time within the target grid is saved as the observation time. If the "valid_time" is configured at the configuration file, the valid_time from the configuration file is saved into the output file. +For MET observation input and CF-compliant NetCDF input with 2D time variable: The latest observation time within the target grid is saved as the observation time. If the "valid_time" is configured in the configuration file, the valid_time from the configuration file is saved into the output file. point2grid Configuration File ----------------------------- @@ -1159,11 +1159,11 @@ _____________________ .. code-block:: none - obs_window = { beg = -5400; end = 5400; } - message_type = []; - obs_quality_inc = []; - obs_quality_exc = []; - version = "VN.N"; + obs_window = { beg = -5400; end = 5400; } + message_type = []; + obs_quality_inc = []; + obs_quality_exc = []; + version = "VN.N"; The configuration options listed above are common to many MET tools and are described in :numref:`config_options`. @@ -1171,7 +1171,7 @@ _____________________ .. code-block:: none - valid_time = "YYYYMMDD_HHMMSS"; + valid_time = "YYYYMMDD_HHMMSS"; This entry is a string to override the observation time into the output and to filter observation data by time. @@ -1179,20 +1179,20 @@ _____________________ .. code-block:: none - var_name_map = [ - { key = "1"; val = "PRES"; }, // GRIB: Pressure - { key = "2"; val = "PRMSL"; }, // GRIB: Pressure reduced to MSL - { key = "7"; val = "HGT"; }, // GRIB: Geopotential height - { key = "11"; val = "TMP"; }, // GRIB: Temperature - { key = "15"; val = "TMAX"; }, // GRIB: Max Temperature - ... - { key = "lat_vname"; val = "NLAT"; }, // NetCDF latitude variable name - { key = "lon_vname"; val = "NLON"; }, // NetCDF longitude varialbe name - ... - ] + var_name_map = [ + { key = "1"; val = "PRES"; }, // GRIB: Pressure + { key = "2"; val = "PRMSL"; }, // GRIB: Pressure reduced to MSL + { key = "7"; val = "HGT"; }, // GRIB: Geopotential height + { key = "11"; val = "TMP"; }, // GRIB: Temperature + { key = "15"; val = "TMAX"; }, // GRIB: Max Temperature + ... + { key = "lat_vname"; val = "NLAT"; }, // NetCDF latitude variable name + { key = "lon_vname"; val = "NLON"; }, // NetCDF longitude variable name + ... + ] This entry is an array of dictionaries, each containing a **GRIB code** string and matching **variable name** string which define a mapping of GRIB code to the output variable names. -The latitude and longitude variables for NetCDF input can be overridden by the configurations. There are two special keys, **lat_vname** and **lon_vname**, are applied to the NetCDF input, not for a GRIB code. +The latitude and longitude variables for NetCDF input can be overridden by the configurations. There are two special keys, **lat_vname** and **lon_vname**, which are applied to the NetCDF input, not for a GRIB code. Point NetCDF to ASCII Python Utility ==================================== @@ -1203,41 +1203,41 @@ The script can be found at: .. code-block:: none - MET_BASE/python/utility/print_pointnc2ascii.py + MET_BASE/python/utility/print_pointnc2ascii.py For how to use the script, issue the command: .. code-block:: none - python3 MET_BASE/python/utility/print_pointnc2ascii.py -h + python3 MET_BASE/python/utility/print_pointnc2ascii.py -h IABP retrieval Python Utilities ==================================== -`International Arctic Buoy Programme (IABP) Data `_ is one of the data types supported by ascii2nc. A utility script that pulls all this data from the web and stores it locally, called get_iabp_from_web.py is included. This script accesses the appropriate webpage and downloads the ascii files for all buoys. It is straightforward, but can be time intensive as the archive of this data is extensive and files are downloaded one at a time. +`International Arctic Buoy Programme (IABP) Data `_ is one of the data types supported by ascii2nc. A utility script that pulls all this data from the web and stores it locally, called get_iabp_from_web.py is included. This script accesses the appropriate webpage and downloads the ASCII files for all buoys. It is straightforward, but can be time intensive as the archive of this data is extensive and files are downloaded one at a time. The script can be found at: .. code-block:: none - MET_BASE/python/utility/get_iabp_from_web.py + MET_BASE/python/utility/get_iabp_from_web.py For how to use the script, issue the command: .. code-block:: none - python3 MET_BASE/python/utility/get_iabp_from_web.py -h + python3 MET_BASE/python/utility/get_iabp_from_web.py -h -Another IABP utility script is included for users, to be run after all files have been downloaded using get_iabp_from_web.py. This script examines all the files and lists those files that contain entries that fall within a user specified range of days. It is called find_iabp_in_timerange.py. +Another IABP utility script is included for users, to be run after all files have been downloaded using get_iabp_from_web.py. This script examines all the files and lists those files that contain entries that fall within a user-specified range of days. It is called find_iabp_in_timerange.py. The script can be found at: .. code-block:: none - MET_BASE/python/utility/find_iabp_in_timerange.py + MET_BASE/python/utility/find_iabp_in_timerange.py For how to use the script, issue the command: .. code-block:: none - python3 MET_BASE/python/utility/find_iabp_in_timerange.py -h + python3 MET_BASE/python/utility/find_iabp_in_timerange.py -h diff --git a/docs/Users_Guide/refs.rst b/docs/Users_Guide/refs.rst index a1411e6d90..609079ee99 100644 --- a/docs/Users_Guide/refs.rst +++ b/docs/Users_Guide/refs.rst @@ -156,7 +156,7 @@ References .. _Ebert-Uphoff-2024: -| Ebert-Uphoff, I.,, 2024: An Investigation of Metrics to Evaluate the Sharpness +| Ebert-Uphoff, I., 2024: An Investigation of Metrics to Evaluate the Sharpness | in AI-Generated Meteorological Imagery. *Draft version - Jan 26, 2024* | @@ -169,8 +169,8 @@ References .. _Efron-2007: -| Efron, B. 2007: Correlation and large-scale significance testing. *Journal* -| of the American Statistical Association*, 102(477), 93-103. +| Efron, B. 2007: Correlation and large-scale significance testing. *Journal of* +| *the American Statistical Association*, 102(477), 93-103. | .. _Epstein-1969: @@ -187,6 +187,14 @@ References | doi: https://doi.org/10.1002/qj.3115 | +.. _Ferro-2011: + +| Ferro, C. A. T., and D. B. Stephenson, 2011: Extremal Dependence Indices: Improved +| Verification Measures for Deterministic Forecasts of Rare Binary Events. +| *Weather and Forecasting*, 26 (5), 699-713. +| doi: https://doi.org/10.1175/WAF-D-10-05030.1 +| + .. _Gilleland-2010: | Gilleland, E., 2010: Confidence intervals for forecast verification. @@ -263,8 +271,8 @@ References .. _Knaff-2003: | Knaff, J.A., M. DeMaria, C.R. Sampson, and J.M. Gross, 2003: Statistical, -| Five-Day Tropical Cyclone Intensity Forecasts Derived from Climatology -| and Persistence. *Weather and Forecasting*, Vol. 18 Issue 2, p. 80-92. +| 5-Day Tropical Cyclone Intensity Forecasts Derived from Climatology +| and Persistence. *Weather and Forecasting*, Vol. 18 Issue 1, p. 80-92. | .. _Mason-2004: @@ -332,7 +340,7 @@ References | Rodwell, M.J., D.S. Richardson, T.D. Hewson and T. Haiden, 2010: A new equitable | score suitable for verifying precipitation in numerical weather prediction. -| *Quarterly Journal of the Royal Meteorological Society*, 136: 1344-1463. +| *Quarterly Journal of the Royal Meteorological Society*, 136: 1344-1363. | https://doi.org/10.1002/qj.656 | diff --git a/docs/Users_Guide/release-notes.rst b/docs/Users_Guide/release-notes.rst index c7b8849683..8958cb5dcb 100644 --- a/docs/Users_Guide/release-notes.rst +++ b/docs/Users_Guide/release-notes.rst @@ -12,105 +12,105 @@ Important issues are listed **in bold** for emphasis. MET Version 13.0.0-beta2 Release Notes (20260508) ------------------------------------------------- - .. dropdown:: Enhancements +.. dropdown:: Enhancements - * **Support @value notation and multiple vertical levels for UGRID NetCDF files** - (`#3254 `_). - * **Enhance Gen-Ens-Prod to support the Ensemble Agreement Scale (EAS) algorithm for calculating probabilities** - (`#3294 `_). - * Refine Regrid-Data-Plane to print warnings about missing fields rather than erroring out - (`#3336 `_). - * Improve the Python embedding handling of gridded data attributes since JSON does not serialize user defined objects like "cartopy.crs.LambertConformal" - (`#3373 `_). + * **Support @value notation and multiple vertical levels for UGRID NetCDF files** + (`#3254 `_). + * **Enhance Gen-Ens-Prod to support the Ensemble Agreement Scale (EAS) algorithm for calculating probabilities** + (`#3294 `_). + * Refine Regrid-Data-Plane to print warnings about missing fields rather than erroring out + (`#3336 `_). + * Improve the Python embedding handling of gridded data attributes since JSON does not serialize user defined objects like "cartopy.crs.LambertConformal" + (`#3373 `_). - .. dropdown:: Bugfixes +.. dropdown:: Bugfixes - * Fix ASCII2NC to handle a wider range of NDBC bad data values - (`#3342 `_). - * Fix support for the "file_type" option in the PCP-Combine "sum" command - (`#3353 `_). - * Fix OpenMP 4.5 compilation errors from the GNU 8.5.0 compiler - (`#3359 `_). - * Fix Grid-Stat to correct the timing information of the SEEPS data written to the NetCDF matched pairs output file - (`#3362 `_). - * Fix TC-RMW to run on BEST tracks (and enhance TC-RMW to more flexibly match track points to gridded data) - (`#3370 `_). - * Fix NetCDF CF convention support to "false_easting" and "false_northing" for lambert conformal projections - (`#3374 `_). + * Fix ASCII2NC to handle a wider range of NDBC bad data values + (`#3342 `_). + * Fix support for the "file_type" option in the PCP-Combine "sum" command + (`#3353 `_). + * Fix OpenMP 4.5 compilation errors from the GNU 8.5.0 compiler + (`#3359 `_). + * Fix Grid-Stat to correct the timing information of the SEEPS data written to the NetCDF matched pairs output file + (`#3362 `_). + * Fix TC-RMW to run on BEST tracks (and enhance TC-RMW to more flexibly match track points to gridded data) + (`#3370 `_). + * Fix NetCDF CF convention support to "false_easting" and "false_northing" for lambert conformal projections + (`#3374 `_). - .. dropdown:: Repository, build, and test +.. dropdown:: Repository, build, and test - * Update MET's development environment to better support RRFS GRIB2 files - (`#3337 `_). - * Update MET's compilation script to support upgraded versions of the dependent libraries - (`#3377 `_). + * Update MET's development environment to better support RRFS GRIB2 files + (`#3337 `_). + * Update MET's compilation script to support upgraded versions of the dependent libraries + (`#3377 `_). - .. dropdown:: METbaseimage testing environment +.. dropdown:: METbaseimage testing environment - * Update METbaseimage to use Python version 3.14 - (`#61 `_). + * Update METbaseimage to use Python version 3.14 + (`#61 `_). - .. note:: +.. note:: - When using the **compile_MET_all.sh** script with this developmental - release, download the appropriate dependency bundle, **tar_files.met-base-develop.tgz**, using - *wget https://dtcenter.ucar.edu/dfiles/code/METplus/MET/installation/tar_files.met-base-develop.tgz* - instead of **tar_files.tgz**. The standard **tar_files.tgz** is intended - for stable releases and may not contain compatible versions of the - required libraries. + When using the **compile_MET_all.sh** script with this developmental + release, download the appropriate dependency bundle, **tar_files.met-base-develop.tgz**, using + *wget https://dtcenter.ucar.edu/dfiles/code/METplus/MET/installation/tar_files.met-base-develop.tgz* + instead of **tar_files.tgz**. The standard **tar_files.tgz** is intended + for stable releases and may not contain compatible versions of the + required libraries. MET Version 13.0.0-beta1 Release Notes (20260205) ------------------------------------------------- - .. dropdown:: Enhancements - - * Minimize the use of temporary files in Stat-Analysis - (`#2698 `_). - * Resolve runtime differences for different GNU/Intel optimization levels for PBL heights in PB2NC - (`#3110 `_). - * **Enhance Grid-Diag to compute mutual information** - (`#3171 `_). - * Enhance MET Python embedding and grid specification strings to support LAEA grids - (`#3230 `_). - * Resolve several SonarQube Reliability issues in MET's develop branch - (`#3253 `_). - * Refine handling of missing data for orographic corrections - (`#3270 `_). - * Enhance the MET tools to return consistent exit codes - (`#3278 `_). - * Enhance the "GEOG_MATCH" interpolation method to print a WARNING about missing topography and land/sea mask inputs - (`#3285 `_). - * Enhance Point2Grid to make the default output value configurable - (`#3297 `_). - * **Refine the logic for setting the default masking "FULL" grid in the MET tools** - (`#3298 `_). - * Enhance PB2NC and IODA2NC to set the "quality_mark_thresh" configuration option as an actual threshold - (`#3307 `_). - - .. dropdown:: Bugfixes - - * Fix the logic to apply "set_attr_grid" before "ShiftRight" - (`#3255 `_). - * Fix ASCII2NC hang when run with an empty input file - (`#3266 `_). - * Fix support for the "set_attr_grid" config option when defining the verification domain - (`#3293 `_). - * Fix dependency checks and compile flag defaults in compile_MET_all.sh - (`#3317 `_). - - .. dropdown:: Repository, build, and test - - * Deprecate and remove the "--enable-mode-graphics" configuration option and corresponding "plot_mode_field" utility - (`#3322 `_). - - .. dropdown:: METbaseimage testing environment - - * Replace deprecated pip install arguments - (`METbaseimage #47 `_). - * Create new hardened and streamlined base image for METviewer - (`METbaseimage #52 `_). +.. dropdown:: Enhancements + + * Minimize the use of temporary files in Stat-Analysis + (`#2698 `_). + * Resolve runtime differences for different GNU/Intel optimization levels for PBL heights in PB2NC + (`#3110 `_). + * **Enhance Grid-Diag to compute mutual information** + (`#3171 `_). + * Enhance MET Python embedding and grid specification strings to support LAEA grids + (`#3230 `_). + * Resolve several SonarQube Reliability issues in MET's develop branch + (`#3253 `_). + * Refine handling of missing data for orographic corrections + (`#3270 `_). + * Enhance the MET tools to return consistent exit codes + (`#3278 `_). + * Enhance the "GEOG_MATCH" interpolation method to print a WARNING about missing topography and land/sea mask inputs + (`#3285 `_). + * Enhance Point2Grid to make the default output value configurable + (`#3297 `_). + * **Refine the logic for setting the default masking "FULL" grid in the MET tools** + (`#3298 `_). + * Enhance PB2NC and IODA2NC to set the "quality_mark_thresh" configuration option as an actual threshold + (`#3307 `_). + +.. dropdown:: Bugfixes + + * Fix the logic to apply "set_attr_grid" before "ShiftRight" + (`#3255 `_). + * Fix ASCII2NC hang when run with an empty input file + (`#3266 `_). + * Fix support for the "set_attr_grid" config option when defining the verification domain + (`#3293 `_). + * Fix dependency checks and compile flag defaults in compile_MET_all.sh + (`#3317 `_). + +.. dropdown:: Repository, build, and test + + * Deprecate and remove the "--enable-mode-graphics" configuration option and corresponding "plot_mode_field" utility + (`#3322 `_). + +.. dropdown:: METbaseimage testing environment + + * Replace deprecated pip install arguments + (`METbaseimage #47 `_). + * Create new hardened and streamlined base image for METviewer + (`METbaseimage #52 `_). MET Upgrade Instructions ======================== @@ -124,109 +124,109 @@ This section summarizes and highlights important changes to MET since version 12 formats. - Changing existing **output data values** generated by MET, typically due to fixing bugs that were computing incorrect output. - - Any other relevant details needed to upgrade from the MET version 12.1.0 to 12.2.0. + - Any other relevant details needed to upgrade from the MET version 12.2.0 to 13.0.0. MET Version 13.0.0 Upgrade Instructions --------------------------------------- .. dropdown:: New or deprecated tools - MET version 13.0.0 adds or removes the following tools: + MET version 13.0.0 adds or removes the following tools: - * The plot_mode_field utility to visualize the NetCDF output generated by the MODE tool - has been deprecated and removed to reduce MET's external library dependencies, including - the Cairo, Freetype, and Pixman libraries and the Ghostscript Fonts package. Users are - encouarged to visualize the NetCDF output from MODE using the plot_data_plane utility - and/or using functionality provided in `METplotpy `_. + * The plot_mode_field utility to visualize the NetCDF output generated by the MODE tool + has been deprecated and removed to reduce MET's external library dependencies, including + the Cairo, Freetype, and Pixman libraries and the Ghostscript Fonts package. Users are + encouraged to visualize the NetCDF output from MODE using the plot_data_plane utility + and/or using functionality provided in `METplotpy `_. .. dropdown:: Command line option changes - MET version 13.0.0 adds, modifies, or removes the following command line options: + MET version 13.0.0 adds, modifies, or removes the following command line options: - * Point2Grid adds a new "-default_value" command line option to override the default - output grid value of bad data (NA or -9999). + * Point2Grid adds a new "-default_value" command line option to override the default + output grid value of bad data (NA or -9999). .. dropdown:: Configuration file changes - MET version 13.0.0 adds, modifies, or removes the following configuration options: + MET version 13.0.0 adds, modifies, or removes the following configuration options: - * The PB2NC and IODA2NC "quality_mark_thresh" configuration file option can now be set as a - threshold string or integer, as before. + * The PB2NC and IODA2NC "quality_mark_thresh" configuration file option can now be set as a + threshold string or integer, as before. - * When processing GDAS surface observations with PB2NC, consider setting - "quality_mark_thresh = <=2||==9;" to retain high quality surface observations - (1 and 2) plus those ignored by the data assimilation system (9). + * When processing GDAS surface observations with PB2NC, consider setting + "quality_mark_thresh = <=2||==9;" to retain high quality surface observations + (1 and 2) plus those ignored by the data assimilation system (9). - * While IODA2NC did previously set "quality_mark_thresh = 0;" in the configuration file, - it was not actually applied by IODA2NC and no quality mark filtering was performed. - Changing the default to "quality_mark_thresh = NA;", which always evaluates to true, - maintains that earlier default behavior. However, changing this setting now applies - the expected filtering logic. + * While IODA2NC did previously set "quality_mark_thresh = 0;" in the configuration file, + it was not actually applied by IODA2NC and no quality mark filtering was performed. + Changing the default to "quality_mark_thresh = NA;", which always evaluates to true, + maintains that earlier default behavior. However, changing this setting now applies + the expected filtering logic. - * Setting the mask grid to "FULL" has been removed from the default configuration files for - Point-Stat, Grid-Stat, Pair-Stat, and Ensemble-Stat to produce more intuitive behavior. - However, if no spatial masking region is requested in the configuration file, then the - "FULL" grid will automatically be added, as noted in :numref:`config_options-mask`. + * Setting the mask grid to "FULL" has been removed from the default configuration files for + Point-Stat, Grid-Stat, Pair-Stat, and Ensemble-Stat to produce more intuitive behavior. + However, if no spatial masking region is requested in the configuration file, then the + "FULL" grid will automatically be added, as noted in :numref:`config_options-mask`. - * Grid-Diag configuration file + * Grid-Diag configuration file - * The "mask.grid" and "mask.poly" entries have changed from strings to arrays of strings - to support the processing of multiple masking regions. + * The "mask.grid" and "mask.poly" entries have changed from strings to arrays of strings + to support the processing of multiple masking regions. - * The "power_spectrum" dictionary is added to configure how missing data values should - handled when computing power spectra. + * The "power_spectrum" dictionary is added to configure how missing data values should + be handled when computing power spectra. - * The new "output_flag" entry is a dictionary specifying the desired output types. + * The new "output_flag" entry is a dictionary specifying the desired output types. - * Gen-Ens-Prod configuration file + * Gen-Ens-Prod configuration file - * The "eas_prob" dictionary is added to configure the Ensemble Agreement Scale (EAS) algorithm. + * The "eas_prob" dictionary is added to configure the Ensemble Agreement Scale (EAS) algorithm. - * The "ensemble_flag.eas" and "ensemble_flag.eas_width" entries are added to enable the - writing of EAS output fields. + * The "ensemble_flag.eas" and "ensemble_flag.eas_width" entries are added to enable the + writing of EAS output fields. - * ConfigConstanst configuration file + * ConfigConstants configuration file - * The "u_wind_field_name", "v_wind_field_name", "wind_speed_field_name", and "wind_direction_field_name" - entries are comma-separated lists of field names to search when read wind data needed to derive winds - or rotate them from being grid-relative to earth-relative. See :numref:`PS_wind_rotation_derivation` - for a description of wind rotation and derivation and :numref:`config_wind_field_names` for the - corresponding configuration options. + * The "u_wind_field_name", "v_wind_field_name", "wind_speed_field_name", and "wind_direction_field_name" + entries are comma-separated lists of field names to search when reading wind data needed to derive winds + or rotate them from being grid-relative to earth-relative. See :numref:`PS_wind_rotation_derivation` + for a description of wind rotation and derivation and :numref:`config_wind_field_names` for the + corresponding configuration options. .. dropdown:: Output format changes - MET version 13.0.0 adds or modifies the following output file formats: + MET version 13.0.0 adds or modifies the following output file formats: - * Grid-Diag output format + * Grid-Diag output format - * The new "mask" dimension, "mask_name" variable, and "mask_size" variable are added - to support the processing of mulitple masking regions. + * The new "mask" dimension, "mask_name" variable, and "mask_size" variable are added + to support the processing of multiple masking regions. - * Existing histogram variables are modified to include the "mask" dimension. + * Existing histogram variables are modified to include the "mask" dimension. - * New information theory variables are added for "entropy", "joint_entropy", and - "mutual_information". + * New information theory variables are added for "entropy", "joint_entropy", and + "mutual_information". - * A new power spectrum "wavenumber" dimension is added, along with "wavenumber" and - "wavelength" variables. "power_spectrum" variables are written for each input, - and "error_power_spectrum" variables contain the differences for multiple inputs. + * A new power spectrum "wavenumber" dimension is added, along with "wavenumber" and + "wavelength" variables. "power_spectrum" variables are written for each input, + and "error_power_spectrum" variables contain the differences for multiple inputs. - * Gen-Ens-Prod output format + * Gen-Ens-Prod output format - * Adds new output variables with names include "EAS" and "EAS_WIDTH" for the Ensemble - Agreement Scale algorithm. + * Adds new output variables with names including "EAS" and "EAS_WIDTH" for the Ensemble + Agreement Scale algorithm. .. dropdown:: Output data changes - MET version 13.0.0 modifies existing output data values in the following ways: + MET version 13.0.0 modifies existing output data values in the following ways: - * Ranked probability statistics found in the RPS line type and written by Ensemble-Stat and the - HiRA algorithm of Point-Stat are modified slightly. As described in MET - `#3396 `_, improvements to the RPS computation - algorithm result in minor changes to the statistics. + * Ranked probability statistics found in the RPS line type and written by Ensemble-Stat and the + HiRA algorithm of Point-Stat are modified slightly. As described in MET + `#3396 `_, improvements to the RPS computation + algorithm result in minor changes to the statistics. .. dropdown:: Additional upgrade instructions - NONE - Recommendations when upgrading to MET version 13.0.0: + Recommendations when upgrading to MET version 13.0.0: - * None + * None diff --git a/docs/Users_Guide/rmw-analysis.rst b/docs/Users_Guide/rmw-analysis.rst index d47adb1a37..939159e8d9 100644 --- a/docs/Users_Guide/rmw-analysis.rst +++ b/docs/Users_Guide/rmw-analysis.rst @@ -65,7 +65,7 @@ ______________________ version = "VN.N"; -The :code:`data` dictionary specifies the name of the 2D and 3D gridded variables to be processed from the TC-RMW NetCDF output files. Its formatting is the same as the :code:`fcst` and :code:`obs` dictionaries, used in the MET statistics tools. However, while the :code:`level` string must be specified, it is not actually used and can remain as its default setting of an empty string. TC-RMW reads data for all variables requested and summarizes it through time for all tracks that meet the filtering criteria described below. +The :code:`data` dictionary specifies the name of the 2D and 3D gridded variables to be processed from the TC-RMW NetCDF output files. Its formatting is the same as the :code:`fcst` and :code:`obs` dictionaries, used in the MET statistics tools. However, while the :code:`level` string must be specified, it is not actually used and can remain as its default setting of an empty string. RMW-Analysis reads data for all variables requested and summarizes it through time for all tracks that meet the filtering criteria described below. The configuration options listed above are common to many MET tools and are described in :numref:`config_options`. @@ -106,7 +106,7 @@ ____________________ The NetCDF output from TC-RMW contains ATCF-formatted storm track information in the :code:`TrackLines` variable. The RMW-Analysis tool parses that track information and applies the filtering criteria listed above. Filtering is only applied for configuration entries that are non-empty lists or strings. The corresponding gridded data is only used for track points that meet all specified filtering criteria. -The :code:`column_thresh_name` and :code:`init_thresh_name` arrays specify the names of ATCF columns whose values should be checked. The former applies to each individual track point while the latter applies to the initial track point (i.e. lead time equals 0). If the filtering criteria is not satisfied for the initial time, the entire track is discarded. Only values from select ATCF columns (LAT, LON, VMAX, MSLP, POUTER, ROUTER, RMW, GUSTS, EYE, DIR, and SPEED) can be thresholded numerically with these options to filter the input data processed. +The :code:`column_thresh_name` and :code:`init_thresh_name` arrays specify the names of ATCF columns whose values should be checked. The former applies to each individual track point while the latter applies to the initial track point (i.e., lead time equals 0). If the filtering criteria is not satisfied for the initial time, the entire track is discarded. Only values from select ATCF columns (LAT, LON, VMAX, MSLP, POUTER, ROUTER, RMW, GUSTS, EYE, DIR, and SPEED) can be thresholded numerically with these options to filter the input data processed. These configuration options are described in :numref:`config_options_tc`. diff --git a/docs/Users_Guide/series-analysis.rst b/docs/Users_Guide/series-analysis.rst index 5804263ff6..0dad453a0e 100644 --- a/docs/Users_Guide/series-analysis.rst +++ b/docs/Users_Guide/series-analysis.rst @@ -12,7 +12,7 @@ The Series-Analysis Tool accumulates statistics separately for each horizontal g Practical Information ===================== -This Series-Analysis tool performs verification of gridded model fields using matching gridded observation fields. It computes a variety of user-selected statistics. These statistics are a subset of those produced by the Grid-Stat tool, with options for statistic types, thresholds, and conditional verification options as discussed in :numref:`grid-stat`. However, these statistics are computed separately for each grid location and accumulated over some series such as time or height, rather than accumulated over the whole domain for a single time or height as is done by Grid-Stat. +The Series-Analysis tool performs verification of gridded model fields using matching gridded observation fields. It computes a variety of user-selected statistics. These statistics are a subset of those produced by the Grid-Stat tool, with options for statistic types, thresholds, and conditional verification options as discussed in :numref:`grid-stat`. However, these statistics are computed separately for each grid location and accumulated over some series such as time or height, rather than accumulated over the whole domain for a single time or height as is done by Grid-Stat. This tool computes statistics for exactly one series each time it is run. Multiple series may be processed by running the tool multiple times. The length of the series to be processed is determined by the first of the following that is greater than one: the number of forecast fields in the configuration file, the number of observation fields in the configuration file, the number of input forecast files, the number of input observation files. Several examples of defining series are described below. @@ -43,8 +43,8 @@ The usage statement for the Series-Analysis tool is shown below: series_analysis has four required arguments and accepts several optional ones. -Required Arguments series_stat -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +Required Arguments for series_analysis +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ 1. The **-fcst file_1 ... file_n | file_list** option specifies the gridded forecast files or ASCII file list of file names to be used, as described in :numref:`ascii_file_lists`. @@ -92,13 +92,13 @@ The Series-Analysis tool produces NetCDF files containing output statistics for .. figure:: figure/series-analysis_Glibert_precip.png - An example of the Gilbert Skill Score for precipitation forecasts at each grid location for a month of files. + An example of the Gilbert Skill Score for precipitation forecasts at each grid location for a month of files. series_analysis Configuration File ---------------------------------- The default configuration file for the Series-Analysis tool named **SeriesAnalysisConfig_default** can be found in the installed *share/met/config* directory. The contents of the configuration file are described in the subsections below. -Note that environment variables may be used when editing configuration files, as described in the :numref:`config_env_vars`. +Note that environment variables may be used when editing configuration files, as described in :numref:`config_env_vars`. ____________________ @@ -176,7 +176,7 @@ The output_stats array controls the type of output that the Series-Analysis tool 7. SL1L2 for Scalar L1L2 Partial Sums (See :numref:`table_PS_format_info_SL1L2`) -8. SAL1L2 for Scalar Anomaly L1L2 Partial Sums climatological data is supplied (See :numref:`table_PS_format_info_SAL1L2`) +8. SAL1L2 for Scalar Anomaly L1L2 Partial Sums when climatological data is supplied (See :numref:`table_PS_format_info_SAL1L2`) 9. PCT for Contingency Table Counts for Probabilistic forecasts (See :numref:`table_PS_format_info_PCT`) @@ -188,4 +188,4 @@ The output_stats array controls the type of output that the Series-Analysis tool 13. GRAD for Gradient Statistics (See :numref:`table_GS_format_info_GRAD`) -.. note:: When the -input option is used, all partial sum and contingency table count columns are required to aggregate statistics across multiple runs. To facilitate this, the output_stats entries for the CTC, SL1L2, SAL1L2, PCT, and GRAD line types can be set to "ALL" to indicate that all available columns for those line types should be written. +.. note:: When the -aggr option is used, all partial sum and contingency table count columns are required to aggregate statistics across multiple runs. To facilitate this, the output_stats entries for the CTC, SL1L2, SAL1L2, PCT, and GRAD line types can be set to "ALL" to indicate that all available columns for those line types should be written. diff --git a/docs/Users_Guide/stat-analysis.rst b/docs/Users_Guide/stat-analysis.rst index 94e23712c6..a1da144e27 100644 --- a/docs/Users_Guide/stat-analysis.rst +++ b/docs/Users_Guide/stat-analysis.rst @@ -9,7 +9,7 @@ Introduction The Stat-Analysis tool ties together results from the Point-Stat, Grid-Stat, Ensemble-Stat, Wavelet-Stat, and TC-Gen tools by providing summary statistical information and a way to filter their STAT output files. It processes the STAT output created by the other MET tools in a variety of ways which are described in this section. -MET version 9.0 adds support for the passing matched pair data (MPR) into Stat-Analysis using a Python script with the "-lookin python ..." option. An example of running Stat-Analysis with Python embedding can be found in :numref:`Appendix F, Section %s `. +MET version 9.0 adds support for passing matched pair data (MPR) into Stat-Analysis using a Python script with the "-lookin python ..." option. An example of running Stat-Analysis with Python embedding can be found in :numref:`Appendix F, Section %s `. Scientific and Statistical Aspects ================================== @@ -29,7 +29,7 @@ The Stat-Analysis "summary" job produces summary information for columns of data Confidence intervals are computed for the mean and standard deviation of the column of data. For the mean, the confidence interval is computed two ways - based on an assumption of normality and also using the bootstrap method. For the standard deviation, the confidence interval is computed using the bootstrap method. In this application of the bootstrap method, the values in the column of data being summarized are resampled, and for each replicated sample, the mean and standard deviation are computed. -The columns to be summarized can be specified in one of two ways. Use the **-line_type** option exactly once to specify a single input line type and use the **-column** option one or more times to select the columns of data to be summarized. Alternatively, use the **-column** option one or more times formatting the entries as **LINE_TYPE:COLUMN**. For example, the RMSE column from the CNT line type can be selected using **-line_type CNT -column RMSE** or using **-column CNT:RMSE**. With the second option, columns from multiple input line types may be selected. For example, **-column CNT:RMSE,CNT:MAE,CTS:CSI** select two CNT columns and one CTS column. +The columns to be summarized can be specified in one of two ways. Use the **-line_type** option exactly once to specify a single input line type and use the **-column** option one or more times to select the columns of data to be summarized. Alternatively, use the **-column** option one or more times formatting the entries as **LINE_TYPE:COLUMN**. For example, the RMSE column from the CNT line type can be selected using **-line_type CNT -column RMSE** or using **-column CNT:RMSE**. With the second option, columns from multiple input line types may be selected. For example, **-column CNT:RMSE,CNT:MAE,CTS:CSI** selects two CNT columns and one CTS column. The WMO mean values are computed in one of three ways, as determined by the configuration file settings for **wmo_sqrt_stats** and **wmo_fisher_stats**. The statistics listed in the first option are square roots. When computing WMO means, the input values are first squared, then averaged, and the square root of the average value is reported. The statistics listed in the second option are correlations to which the Fisher transformation is applied. For any statistic not listed, the WMO mean is computed as a simple arithmetic mean. The **WMO_TYPE** output column indicates the method applied (**SQRT, FISHER**, or **MEAN**). The **WMO_MEAN** and **WMO_WEIGHTED_MEAN** columns contain the unweighted and weighted means, respectively. The value listed in the **TOTAL** column of each input line is used as the weight. @@ -45,9 +45,9 @@ The Stat-Analysis "aggregate" job aggregates values from multiple STAT lines of Aggregate STAT Lines and Produce Aggregated Statistics ------------------------------------------------------ -The Stat-Analysis "aggregate-stat" job aggregates multiple STAT lines of the same type together and produces relevant statistics from the aggregated line. This may be done in the same manner listed above in :numref:`StA_Aggregated-values-from`. However, rather than writing out the aggregated STAT line itself, the relevant statistics generated from that aggregated line are provided in the output. Specifically, if a contingency table line type (FHO, CTC, PCT, MCTC, or NBRCTC) has been aggregated, contingency table statistics (CTS, ECLV, PSTD, MCTS, or NBRCTS) line types can be computed. If a partial sums line type (SL1L2 or SAL1L2) has been aggregated, the continuous statistics (CNT) line type can be computed. If a vector partial sums line type (VL1L2 or VAL1L2) has been aggregated, the vector continuous statistics (VCNT) line type can be computed. For ensembles, the ORANK line type can be accumulated into ECNT, RPS, RHIST, PHIST, RELP, or SSVAR output. If a SEEPS matched pair line type (SEEPS_MPR) has been aggregated, the aggregated SEEPS line type (SEEPS) can be computed. If the matched pair line type (MPR) has been aggregated, may output line types (FHO, CTC, CTS, CNT, MCTC, MCTS, SL1L2, SAL1L2, VL1L2, VCNT, WDIR, PCT, PSTD, PJC, PRC, or ECLV) can be computed. Multiple output line types may be specified for each "aggregate-stat" job, as long as each output is derivable from the input. +The Stat-Analysis "aggregate-stat" job aggregates multiple STAT lines of the same type together and produces relevant statistics from the aggregated line. This may be done in the same manner listed above in :numref:`StA_Aggregated-values-from`. However, rather than writing out the aggregated STAT line itself, the relevant statistics generated from that aggregated line are provided in the output. Specifically, if a contingency table line type (FHO, CTC, PCT, MCTC, or NBRCTC) has been aggregated, contingency table statistics (CTS, ECLV, PSTD, MCTS, or NBRCTS) line types can be computed. If a partial sums line type (SL1L2 or SAL1L2) has been aggregated, the continuous statistics (CNT) line type can be computed. If a vector partial sums line type (VL1L2 or VAL1L2) has been aggregated, the vector continuous statistics (VCNT) line type can be computed. For ensembles, the ORANK line type can be accumulated into ECNT, RPS, RHIST, PHIST, RELP, or SSVAR output. If a SEEPS matched pair line type (SEEPS_MPR) has been aggregated, the aggregated SEEPS line type (SEEPS) can be computed. If the matched pair line type (MPR) has been aggregated, many output line types (FHO, CTC, CTS, CNT, MCTC, MCTS, SL1L2, SAL1L2, VL1L2, VCNT, WDIR, PCT, PSTD, PJC, PRC, or ECLV) can be computed. Multiple output line types may be specified for each "aggregate-stat" job, as long as each output is derivable from the input. -When aggregating the matched pair line type (MPR), additional required job command options are determined by the requested output line type(s). For example, the "-out_thresh" (or "-out_fcst_thresh" and "-out_obs_thresh" options) are required to compute contingnecy table counts (FHO, CTC) or statistics (CTS). Those same job command options can also specify filtering thresholds when computing continuous partial sums (SL1L2, SAL1L2) or statistics (CNT). Output is written for each threshold specified. +When aggregating the matched pair line type (MPR), additional required job command options are determined by the requested output line type(s). For example, the "-out_thresh" (or "-out_fcst_thresh" and "-out_obs_thresh" options) are required to compute contingency table counts (FHO, CTC) or statistics (CTS). Those same job command options can also specify filtering thresholds when computing continuous partial sums (SL1L2, SAL1L2) or statistics (CNT). Output is written for each threshold specified. When aggregating the matched pair line type (MPR) and computing an output contingency table statistics (CTS) or continuous statistics (CNT) line type, the bootstrapping method can be applied to compute confidence intervals. The bootstrapping method is applied here in the same way that it is applied in the statistics tools. For a set of n matched forecast-observation pairs, the matched pairs are resampled with replacement many times. For each replicated sample, the corresponding statistics are computed. The confidence intervals are derived from the statistics computed for each replicated sample. @@ -60,17 +60,17 @@ The Stat-Analysis "ss_index", "go_index", and "cbs_index" jobs calculate skill s In general, a skill score index is computed over several terms and the number and definition of those terms is configurable. It is computed by aggregating the output from earlier runs of the Point-Stat and/or Grid-Stat tools over one or more cases. When configuring a skill score index job, the following requirements apply: -1. Exactly two models names must be chosen. The first is interpreted as the forecast model and the second is the reference model, against which the performance of the forecast should be measured. Specify this with the "model" configuration file entry or using the "-model" job command option. +1. Exactly two model names must be chosen. The first is interpreted as the forecast model and the second is the reference model, against which the performance of the forecast should be measured. Specify this with the "model" configuration file entry or using the "-model" job command option. 2. The forecast variable name, level, lead time, line type, column, and weight options must be specified. If the value remains constant for all the terms, set it to an array of length one. If the value changes for at least one term, specify an array entry for each term. Specify these with the "fcst_var", "fcst_lev", "lead_time", "line_type", "column", and "weight" configuration file entries, respectively, or use the corresponding job command options. -3. While these line types are required, additional options may be provided for each term, including the observation type ("obtype"), verification region ("vx_mask"), and interpolation method ("interp_mthd"). Specify each as single value or provide a value for each term. +3. While these line types are required, additional options may be provided for each term, including the observation type ("obtype"), verification region ("vx_mask"), and interpolation method ("interp_mthd"). Specify each as a single value or provide a value for each term. 4. Only the SL1L2 and CTC input line types are supported, and the input Point-Stat and/or Grid-Stat output must contain these line types. -5. For the SL1L2 line type, set the "column" entry to the CNT output column that contains the statistic of interest (e.g. RMSE for root-mean-squared-error). Note, only those continuous statistics that are derivable from SL1L2 lines can be used. +5. For the SL1L2 line type, set the "column" entry to the CNT output column that contains the statistic of interest (e.g., RMSE for root-mean-squared-error). Note, only those continuous statistics that are derivable from SL1L2 lines can be used. -6. For the CTC line type, set the "column" entry to the CTS output column that contains the statistic of intereest (e.g. PODY for probability of detecting yes). Note, consider specifying the "fcst_thresh" for the CTC line type. +6. For the CTC line type, set the "column" entry to the CTS output column that contains the statistic of interest (e.g., PODY for probability of detecting yes). Note, consider specifying the "fcst_thresh" for the CTC line type. For each term, all matching SL1L2 (or CTC) input lines are aggregated separately for the forecast and reference models. The requested statistic ("column") is derived from the aggregated partial sums or counts. For each term, a skill score is defined as: @@ -78,7 +78,7 @@ For each term, all matching SL1L2 (or CTC) input lines are aggregated separately Where :math:`s_{fcst}` and :math:`s_{ref}` are the aggregated forecast and reference statistics, respectively. Next, a weighted average is computed from the skill scores for each term: - .. math:: ss_{avg} = \frac{1}{n} \sum_{i=1}^{n} w_i * ss_i +.. math:: ss_{avg} = \frac{1}{n} \sum_{i=1}^{n} w_i * ss_i Where, :math:`w_i` and :math:`ss_i` are the weight and skill score for each term and :math:`n` is the number of terms. Finally, the skill score index is computed as: @@ -240,7 +240,7 @@ The Stat-Analysis "ramp" job identifies ramp events (large increases or decrease Wind Direction Statistics ------------------------- -The Stat-Analysis "aggregate_stat" job can read vector partial sums and derive wind direction error statistics (WDIR). The vector partial sums (VL1L2 or VAL1L2) or matched pairs (MPR) for the UGRD and VGRD must have been computed in a previous step, i.e. by Point-Stat or Grid-Stat tools. This job computes an average forecast wind direction and an average observed wind direction along with their difference. The output is in degrees. In Point-Stat and Grid-Stat, the UGRD and VGRD can be verified using thresholds on their values or on the calculated wind speed. If thresholds have been applied, the wind direction statistics are calculated for each threshold. +The Stat-Analysis "aggregate_stat" job can read vector partial sums and derive wind direction error statistics (WDIR). The vector partial sums (VL1L2 or VAL1L2) or matched pairs (MPR) for the UGRD and VGRD must have been computed in a previous step, i.e., by Point-Stat or Grid-Stat tools. This job computes an average forecast wind direction and an average observed wind direction along with their difference. The output is in degrees. In Point-Stat and Grid-Stat, the UGRD and VGRD can be verified using thresholds on their values or on the calculated wind speed. If thresholds have been applied, the wind direction statistics are calculated for each threshold. The first step in verifying wind direction is running the Grid-Stat and/or Point-Stat tools to verify each forecast of interest and generate the VL1L2 or MPR line(s). When running these tools, please note: @@ -258,7 +258,7 @@ Once the appropriate lines have been generated for each verification time of int 2. For the "AGGR_WDIR" line, the input VL1L2 lines are first aggregated into a single line of partial sums where the weight for each line is determined by the number of points it represents. From this aggregated line, the mean forecast wind direction, observation wind direction, and the associated error are computed and written out. -Note that wind direction is undefined for zero length vectors. Any inputs with zero length forecast or observation vectors are automatically excluded from the "ROW_MEAN_WDIR" output line since the direction error is also undefined. However, those inputs are included in the aggregated statistics in the "AGGR_WDIR" output line. For this reason, the "TOTAL" column value, indicating the number of inputs used, may larger in "AGGR_WDIR" than "ROW_MEAN_WDIR", from which zero length vectors have been excluded. Users can also specify the "-out_wind_thresh gt0 -out_wind_logic INTERSECTION" job command options to explicitly exclude all pairs containing zero length vectors from the analysis. In that case, the "ROW_MEAN_WDIR" and "AGGR_WDIR" "TOTAL" column values would match. +Note that wind direction is undefined for zero length vectors. Any inputs with zero length forecast or observation vectors are automatically excluded from the "ROW_MEAN_WDIR" output line since the direction error is also undefined. However, those inputs are included in the aggregated statistics in the "AGGR_WDIR" output line. For this reason, the "TOTAL" column value, indicating the number of inputs used, may be larger in "AGGR_WDIR" than "ROW_MEAN_WDIR", from which zero length vectors have been excluded. Users can also specify the "-out_wind_thresh gt0 -out_wind_logic INTERSECTION" job command options to explicitly exclude all pairs containing zero length vectors from the analysis. In that case, the "ROW_MEAN_WDIR" and "AGGR_WDIR" "TOTAL" column values would match. Practical Information ===================== @@ -284,7 +284,7 @@ The usage statement for the Stat-Analysis tool is shown below: stat_analysis has two required arguments and accepts several optional ones. -In the usage statement for the Stat-Analysis tool, some additional terminology is introduced. In the Stat-Analysis tool, the term "job" refers to a set of tasks to be performed after applying user-specified options (i.e., "filters"). The filters are used to pare down a collection of output from the MET statistics tools to only those lines that are desired for the analysis. The job and its filters together comprise the "job command line". The "job command line" may be specified either on the command line to run a single analysis job or within the configuration file to run multiple analysis jobs at the same time. If jobs are specified in both the configuration file and the command line, only the jobs indicated in the configuration file will be run. The various jobs types are described in :numref:`Des_components_STAT_analysis_tool` and the filtering options are described in :numref:`stat_analysis-configuration-file`. +In the usage statement for the Stat-Analysis tool, some additional terminology is introduced. In the Stat-Analysis tool, the term "job" refers to a set of tasks to be performed after applying user-specified options (i.e., "filters"). The filters are used to pare down a collection of output from the MET statistics tools to only those lines that are desired for the analysis. The job and its filters together comprise the "job command line". The "job command line" may be specified either on the command line to run a single analysis job or within the configuration file to run multiple analysis jobs at the same time. If jobs are specified in both the configuration file and the command line, only the jobs indicated in the configuration file will be run. The various job types are described in :numref:`Des_components_STAT_analysis_tool` and the filtering options are described in :numref:`stat_analysis-configuration-file`. Required Arguments for stat_analysis ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ @@ -322,7 +322,7 @@ stat_analysis Configuration File The default configuration file for the Stat-Analysis tool named **STATAnalysisConfig_default** can be found in the installed *share/met/config* directory. The version used for the example run in :numref:`installation` is also available in *scripts/config*. Like the other configuration files described in this document, it is recommended that users make a copy of these files prior to modifying their contents. -The configuration file for the Stat-Analysis tool is optional. Users may find it more convenient initially to run Stat-Analysis jobs on the command line specifying job command options directly. Once the user has a set of or more jobs they would like to run routinely on the output of the MET statistics tools, they may find grouping those jobs together into a configuration file to be more convenient. +The configuration file for the Stat-Analysis tool is optional. Users may find it more convenient initially to run Stat-Analysis jobs on the command line specifying job command options directly. Once the user has a set of one or more jobs they would like to run routinely on the output of the MET statistics tools, they may find grouping those jobs together into a configuration file to be more convenient. Most of the user-specified parameters listed in the Stat-Analysis configuration file are used to filter the ASCII statistical output from the MET statistics tools down to a desired subset of lines over which statistics are to be computed. Only output that meets all of the parameters specified in the Stat-Analysis configuration file will be retained. @@ -330,7 +330,7 @@ In earlier versions, the Stat-Analysis tool always performed a two step process In practice, the benefit of writing a temporary file varies considerably based on the amount of input data and filtering options. When running multiple jobs on a small subset of the input data, defining common filtering criteria in the top-level section should enable Stat-Analysis to run more efficiently. Applying the common filters once is faster than re-applying them for multiple jobs. However, if running a single job or with no common filtering criteria, writing an intermediate temporary file would provide no benefit. -Filtering options specified in the first section of the configuration file are applied to every task in the **jobs** entry. However, if an individual job specifies an option already set above, the job-specific setting overrides the top-level one. For example, if **model = ["Run1, "Run2"];** is set at the top but a job sets **-model Run1**, that job will be performed only on the "Run1" data. Also note that environment variables may be used when editing configuration files, as described in the :numref:`pb2nc configuration file` for the PB2NC tool. +Filtering options specified in the first section of the configuration file are applied to every task in the **jobs** entry. However, if an individual job specifies an option already set above, the job-specific setting overrides the top-level one. For example, if **model = ["Run1", "Run2"];** is set at the top but a job sets **-model Run1**, that job will be performed only on the "Run1" data. Also note that environment variables may be used when editing configuration files, as described in :numref:`pb2nc configuration file` for the PB2NC tool. ________________________ @@ -408,7 +408,7 @@ ___________________ obs_init_exc = []; obs_init_hour = []; -These time filtering options are the same as described above but applied to initialization times rather than valid times. These selections may be further refined by using the **"-fcst_init_beg", "-fcst_init_end", "-fcst_init_inc", "-fcst_init_exc", "-fcst_init_hour"," "-obs_init_beg", "-obs_init_end", "-obs_init_inc", "-obs_init_exc"** and **"-obs_init_hour"** options within the job command line. +These time filtering options are the same as described above but applied to initialization times rather than valid times. These selections may be further refined by using the **"-fcst_init_beg", "-fcst_init_end", "-fcst_init_inc", "-fcst_init_exc", "-fcst_init_hour", "-obs_init_beg", "-obs_init_end", "-obs_init_inc", "-obs_init_exc"** and **"-obs_init_hour"** options within the job command line. ___________________ @@ -485,7 +485,7 @@ ___________________ alpha = []; -The user may specify a comma-separated list alpha confidence values to be used for all analyses. If alpha values are listed, the analyses will be performed on their union. These selections may be further refined by using the **"-alpha"** option within the job command line. +The user may specify a comma-separated list of alpha confidence values to be used for all analyses. If alpha values are listed, the analyses will be performed on their union. These selections may be further refined by using the **"-alpha"** option within the job command line. ___________________ @@ -511,7 +511,7 @@ ___________________ ss_index_name = "SS_INDEX"; ss_index_vld_thresh = 1.0; -The ss_index_name and ss_index_vld_thresh options are used to define a skill score index. The ss_index_name entry is a string which defines the output name for the current skill score index configuration. The ss_index_vld_thresh entry is a number between 0.0 and 1.0 that defines the required ratio of valid terms. If the ratio of valid skill score index terms to the total is less than than this number, no output is written for that case. The default value of 1.0 indicates that all terms are required. +The ss_index_name and ss_index_vld_thresh options are used to define a skill score index. The ss_index_name entry is a string which defines the output name for the current skill score index configuration. The ss_index_vld_thresh entry is a number between 0.0 and 1.0 that defines the required ratio of valid terms. If the ratio of valid skill score index terms to the total is less than this number, no output is written for that case. The default value of 1.0 indicates that all terms are required. ___________________ @@ -528,11 +528,11 @@ The user may specify one or more analysis jobs to be performed on the STAT lines All possible tasks for **job_name** are listed in :numref:`Des_components_STAT_analysis_tool`. .. role:: raw-html(raw) - :format: html + :format: html .. _Des_components_STAT_analysis_tool: -.. list-table:: Description of components of the job command lines for the Stat-Analysis tool.Variables, levels, and weights used to compute the GO Index. +.. list-table:: Description of components of the job command lines for the Stat-Analysis tool. :widths: 15 55 20 :header-rows: 1 @@ -614,7 +614,7 @@ This job command option is extremely useful. It can be used multiple times to sp -column_str col_name string -column_str_exc col_name string -The column filtering options may be used when the **-line_type** has been set to a single value. These options take two arguments, the name of the data column to be used followed by a value, string, or threshold to be applied. If multiple column_min/max/eq/thresh/str options are listed, the job will be performed on their intersection. Each input line is only retained if its value meets the numeric filtering criteria defined, matches one of the strings defined by the **-column_str** option, or does not match any of the string defined by the **-column_str_exc** option. Multiple filtering strings may be listed using commas. Defining thresholds in MET is described in :numref:`config_options`. +The column filtering options may be used when the **-line_type** has been set to a single value. These options take two arguments, the name of the data column to be used followed by a value, string, or threshold to be applied. If multiple column_min/max/eq/thresh/str options are listed, the job will be performed on their intersection. Each input line is only retained if its value meets the numeric filtering criteria defined, matches one of the strings defined by the **-column_str** option, or does not match any of the strings defined by the **-column_str_exc** option. Multiple filtering strings may be listed using commas. Defining thresholds in MET is described in :numref:`config_options`. .. code-block:: none @@ -652,7 +652,7 @@ The example above reads MPR lines, stratifies the data by forecast variable name -mask_poly file -mask_sid file|list -When processing input MPR lines, these options may be used to define a masking grid, polyline, or list of station ID's to filter the matched pair data geographically prior to computing statistics. The **-mask_sid** option is a station ID masking file or a comma-separated list of station ID's for filtering the matched pairs spatially. See the description of the "sid" entry in :numref:`config_options`. +When processing input MPR lines, these options may be used to define a masking grid, polyline, or list of station IDs to filter the matched pair data geographically prior to computing statistics. The **-mask_sid** option is a station ID masking file or a comma-separated list of station IDs for filtering the matched pairs spatially. See the description of the "sid" entry in :numref:`config_options`. .. code-block:: none @@ -782,7 +782,7 @@ Job: ss_index, go_index, cbs_index While the inputs for the "ss_index", "go_index", and "cbs_index" jobs may vary, the output is the same. By default, the job output is written to the screen or to a "-out" file, if specified. If the "-out_stat" job command option is specified, a STAT output file is written containing the skill score index (SSIDX) output line type. -The SSIDX line type consists of the common STAT header columns described in :numref:`table_PS_header_info_point-stat_out` followed by the columns described below. In general, when multiple input header strings are encountered, the output is reported as a comma-separated list of the unique values. The "-set_hdr" job command option can be used to override any of the output header strings (e.g. "-set_hdr VX_MASK MANY" sets the output VX_MASK column to "MANY"). Special logic applied to some of the STAT header columns are also described below. +The SSIDX line type consists of the common STAT header columns described in :numref:`table_PS_header_info_point-stat_out` followed by the columns described below. In general, when multiple input header strings are encountered, the output is reported as a comma-separated list of the unique values. The "-set_hdr" job command option can be used to override any of the output header strings (e.g., "-set_hdr VX_MASK MANY" sets the output VX_MASK column to "MANY"). Special logic applied to some of the STAT header columns are also described below. .. _table_SA_format_info_SSIDX: diff --git a/docs/Users_Guide/tc-diag.rst b/docs/Users_Guide/tc-diag.rst index 34a92d5669..74c6f06595 100644 --- a/docs/Users_Guide/tc-diag.rst +++ b/docs/Users_Guide/tc-diag.rst @@ -13,7 +13,7 @@ Originally developed for the Statistical Hurricane Intensity Prediction Scheme ( .. note:: A future version of the tool will include the capability to remove the model's own vortex, which will allow the user to specify any arbitrary track (such as the operational center's official forecast). Until then, users are advised that the track selected must be consistent with the model's predicted track. -TC-Diag is run once for each initialization time to produce diagnostics for each user-specified combination of TC tracks and model fields. The user provides track data (such as one or more ATCF a-deck track files), along with track filtering criteria as needed, to select one or more tracks to be processed. The user also provides gridded model data from which diagnostics should be computed. Gridded data can be provided for multiple concurrent storms, multiple models, and/or multiple domains (i.e. parent and nest) in a single run. +TC-Diag is run once for each initialization time to produce diagnostics for each user-specified combination of TC tracks and model fields. The user provides track data (such as one or more ATCF a-deck track files), along with track filtering criteria as needed, to select one or more tracks to be processed. The user also provides gridded model data from which diagnostics should be computed. Gridded data can be provided for multiple concurrent storms, multiple models, and/or multiple domains (i.e., parent and nest) in a single run. TC-Diag first determines the list of valid times that appear in any one of the tracks. For each valid time, it processes all track points for that time. For each track point, it reads the gridded model fields requested in the configuration file and transforms the gridded data to a range-azimuth cylindrical coordinates grid, as described for the TC-RMW tool in :numref:`tc-rmw`. For each domain, it writes the range-azimuth data to a temporary NetCDF file, as described in :numref:`Contributor's Guide Section %s `. @@ -39,7 +39,7 @@ The following sections describe the usage statement, required arguments, and opt Usage: tc_diag -data domain tech_id_list [ file_1 ... file_n | file_list ] - -deck file + -deck path -config file [-outdir path] [-log file] @@ -52,7 +52,7 @@ Required Arguments for tc_diag 1. The **-data domain tech_id_list [ file_1 ... file_n | file_list ]** option specifies a domain name, a comma-separated list of ATCF tech ID's, and a list of gridded data files or an ASCII file containing a list of files to be used, as described in :numref:`ascii_file_lists`. Specify **-data** once for each gridded data source. -2. The **-deck source** option is the ATCF format track data source. +2. The **-deck path** option is the ATCF format track data source. 3. The **-config file** option is the TCDiagConfig file to be used. The contents of the configuration file are discussed below. @@ -126,17 +126,17 @@ Configuring Domain Information } ]; -The **domain_info** entry is an array of dictionaries. Each dictionary consists of five entries. The **domain** entry is a user-specified string that provides a name for the domain. Each **domain** name must also appear in a **-deck** command line option, and the reverse is also true. +The **domain_info** entry is an array of dictionaries. Each dictionary consists of five entries. The **domain** entry is a user-specified string that provides a name for the domain. Each **domain** name must also appear in a **-data** command line option, and the reverse is also true. The **n_range** entry is an integer specifying the number of equally spaced range intervals in the range-azimuth grid to be used for this data source. -The **n_azimuth** entry is an integer specifying the number of equally spaced azimuth intervals in the range-azimuth grid to be used for this data source. The azimuthal grid spacing is 360 / **n_azimuth** degrees. Azimuths are defined by MET as *degrees clockwise* from due east. However, the TC-Diag Python code expects them as *radians counter-clockwise* from due east. The **tc_diag_driver/post_resample_driver.py** driver script performs the neccessary rotation and conversion operations. +The **n_azimuth** entry is an integer specifying the number of equally spaced azimuth intervals in the range-azimuth grid to be used for this data source. The azimuthal grid spacing is 360 / **n_azimuth** degrees. Azimuths are defined by MET as *degrees clockwise* from due east. However, the TC-Diag Python code expects them as *radians counter-clockwise* from due east. The **tc_diag_driver/post_resample_driver.py** driver script performs the necessary rotation and conversion operations. The **delta_range_km** entry is a floating point value specifying the spacing of the range rings in kilometers. The **diag_script** entry is an array of strings. Each string specifies the path to a Python script to be executed to compute diagnostics from the transformed cylindrical coordinates data for this domain. When multiple Python diagnostics scripts are run, the union of the diagnostics computed are written to the output. -The **override_diags** entry is an array of strings. Each string specifies the name of diagnostic value to be used for that domain. If set to an empty list, all diagnostics computed by the Python scripts in **diag_script** for that domain will be used. If non-empty, only the specific diagnostics listed will be used. +The **override_diags** entry is an array of strings. Each string specifies the name of a diagnostic value to be used for that domain. If set to an empty list, all diagnostics computed by the Python scripts in **diag_script** for that domain will be used. If non-empty, only the specific diagnostics listed will be used. In the default configuration, seen above, the same Python script is run for both the *parent* and *nest* domains, each using a different configuration file. For the *parent* domain, all computed diagnostics are used since **override_diags** is empty. For the *nest* domain, only the specific diagnostics listed in **override_diags** are used to override the *parent* values. In general, diagnostics computed earlier in the list of **domain_info** entries can be overridden by diagnostics computed later in the list. @@ -163,7 +163,7 @@ Configuring regridding options shape = SQUARE; } -The **regrid** dictionary is common to multiple MET tools and is described in :numref:`config_options`. It specifies how the input data should be regridded to cylindrical coordinates prior to compute diagnostics. It can be specified separately in each **data.field** array entry, described below. The default setting uses bilinear interpolation for all fields. +The **regrid** dictionary is common to multiple MET tools and is described in :numref:`config_options`. It specifies how the input data should be regridded to cylindrical coordinates prior to computing diagnostics. It can be specified separately in each **data.field** array entry, described below. The default setting uses bilinear interpolation for all fields. Configuring Fields, Levels, and Domains ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ @@ -207,7 +207,7 @@ Configuring Vortex Removal Option .. code-block:: none - vortex_removel = FALSE; + vortex_removal = FALSE; The **vortex_removal** flag entry is a boolean specifying whether or not vortex removal logic should be applied. @@ -233,7 +233,7 @@ If true, all input fields are read efficiently from each file in a single call. These three flag entries are booleans specifying what output data types should be written. At least one of these flags must be set to true. - The **nc_cyl_grid_flag** entry controls the writing of a NetCDF file containing the cylindrical coordinate range-azimuth data used to compute the diagnostics. These files are written with a `_cyl_grid_{domain}.nc` suffix, where `{domain}` is the domain name specified in the configuration file. One output file is written for each combination of model track and domain. - - The **nc_diag_file** entry controls the writing of the computed diagnostics to a NetCDF file. These files are written with a `_diag.nc` suffix. One output file is written for each model track processed. + - The **nc_diag_flag** entry controls the writing of the computed diagnostics to a NetCDF file. These files are written with a `_diag.nc` suffix. One output file is written for each model track processed. - The **cira_diag_flag** entry controls the writing of the computed diagnostics to a formatted ASCII output file. These files are written with a `_diag.dat` suffix. One output file is written for each model track processed. .. code-block:: none @@ -257,15 +257,15 @@ These output files contain tabular ASCII data with diagnostic values either extr - The **STORM DATA** section contains single diagnostic values either extracted from the ATCF track file or computed from the cylindrical grid for each forecast lead time. This section begins with a line named **TIME** defining the forecast lead time of each track point in hours. The following lines contain the requested storm diagnostics. For example, **MAXWIND** contains the maximum wind speed reported in the ATCF track file and **SST** contains the average sea surface temperature computed in the range/azimuth grid. - - The **SOUNDING DATA** section contains diagnostics computed separately for each vertical level. The vertical levels are typically the surface (e.g. **SURF**) followed by pressure levels (e.g. **0850**). This section begins with two lines named **NLEV** and **TIME** defining the number of vertical levels and their values and the forecast lead times for which diagnostics were computed, respectively. The level name is appended to each diagnostic name. For example, the **T_0850** contains the average temperature value within the range/azimuth grid at the 850 mb pressure level. + - The **SOUNDING DATA** section contains diagnostics computed separately for each vertical level. The vertical levels are typically the surface (e.g., **SURF**) followed by pressure levels (e.g., **0850**). This section begins with two lines named **NLEV** and **TIME** defining the number of vertical levels and their values and the forecast lead times for which diagnostics were computed, respectively. The level name is appended to each diagnostic name. For example, the **T_0850** contains the average temperature value within the range/azimuth grid at the 850 mb pressure level. - Each diagnostic output line contains: - - Diagnostic name (with or without the level) e.g. **SHR_MAG** for magnitude of wind shear + - Diagnostic name (with or without the level) e.g., **SHR_MAG** for magnitude of wind shear - - Units string enclosed in parenthesis e.g. **(KT)** for knots + - Units string enclosed in parentheses e.g., **(KT)** for knots - - The diagnostic values computed for each lead time e.g. **13 10 14 ...** + - The diagnostic values computed for each lead time e.g., **13 10 14 ...** **NetCDF Diagnostics Output** @@ -285,7 +285,7 @@ When the **nc_diag_flag** configuration entry is set to true, a NetCDF output fi - Vertical dimension for the number of pressure levels .. role:: raw-html(raw) - :format: html + :format: html .. _table_TC-Diag_Variables_NetCDF_diagnostics: @@ -340,11 +340,11 @@ When the **nc_diag_flag** configuration entry is set to true, a NetCDF output fi **NetCDF Range-Azimuth Output** -When the **nc_rng_azi_flag** configuration entry is set to true, a NetCDF output file containing the cylindrical range-azimuth data is written for each combination of model track provided and domain specified. For example, if three model tracks are provided and data for both *parent* and *nest* domains are provided, six of these NetCDF output files will be written. +When the **nc_cyl_grid_flag** configuration entry is set to true, a NetCDF output file containing the cylindrical range-azimuth data is written for each combination of model track provided and domain specified. For example, if three model tracks are provided and data for both *parent* and *nest* domains are provided, six of these NetCDF output files will be written. The NetCDF range-azimuth output is named using the **output_base_format**, described above, followed by **_cyl_grid_{DOMAIN}.nc**, where **{DOMAIN}** is specified by the **domain** string in each **domain_info** array entry. -This NetCDF file contains a concatenation of the data from the temporary NetCDF files created for each track point. For each track point, TC-Diag creates a temporary NetCDF file and calls Python code to read the cylindrical grid data and compute diagnostics. By default, these temporary NetCDF files are deleted at the end of each run, but if the **nc_rng_azi_flag** is true, the data for each track point is concatenated into a single output file for each track. +This NetCDF file contains a concatenation of the data from the temporary NetCDF files created for each track point. For each track point, TC-Diag creates a temporary NetCDF file and calls Python code to read the cylindrical grid data and compute diagnostics. By default, these temporary NetCDF files are deleted at the end of each run, but if the **nc_cyl_grid_flag** is true, the data for each track point is concatenated into a single output file for each track. .. note:: Setting the **MET_KEEP_TEMP_FILE** (:numref:`met_keep_temp_file`) environment variable retains the temporary NetCDF cylindrical grid files for development, testing, and debugging purposes. @@ -370,7 +370,7 @@ The NetCDF range-azimuth file contains the dimensions and variables shown in :nu - Vertical dimension for the number of pressure levels .. role:: raw-html(raw) - :format: html + :format: html .. _table_TC-Diag_Variables_NetCDF_range_azimuth: @@ -455,12 +455,12 @@ The NetCDF range-azimuth file contains the dimensions and variables shown in :nu - Longitude in degrees east for each range-azimuth grid point - Double * - single level data - (e.g. TMP_Z2, PRMSL_L0) + (e.g., TMP_Z2, PRMSL_L0) - time, range, azimuth - Gridded range-azimuth data on a single level - Double * - pressure level data - (e.g. TMP, HGT) + (e.g., TMP, HGT) - time, pressure, range, azimuth - Gridded range-azimuth data on pressure levels - Double diff --git a/docs/Users_Guide/tc-gen.rst b/docs/Users_Guide/tc-gen.rst index c203ea7f1e..9a646f7804 100644 --- a/docs/Users_Guide/tc-gen.rst +++ b/docs/Users_Guide/tc-gen.rst @@ -7,7 +7,7 @@ TC-Gen Tool Introduction ============ -The TC-Gen tool provides verification of deterministic and probabilistic tropical cyclone genesis forecasts in the ATCF file and shapefile formats. Producing reliable tropical cyclone genesis forecasts is an important metric for global numerical weather prediction models. This tool ingests deterministic model output post-processed by genesis tracking software (e.g. GFDL vortex tracker), ATCF edeck files containing probability of genesis forecasts, operational shapefile warning areas, and ATCF reference track dataset(s) (e.g. Best Track analysis and CARQ operational tracks). It writes categorical counts and statistics. The capability to modify the spatial and temporal tolerances when matching forecasts to reference genesis events, as well as scoring those matched pairs, gives users the ability to condition the criteria based on model performance and/or conduct sensitivity analyses. Statistical aspects are outlined in :numref:`tc-gen_stat_aspects` and practical aspects of the TC-Gen tool are described in :numref:`tc-gen_practical_info`. +The TC-Gen tool provides verification of deterministic and probabilistic tropical cyclone genesis forecasts in the ATCF file and shapefile formats. Producing reliable tropical cyclone genesis forecasts is an important metric for global numerical weather prediction models. This tool ingests deterministic model output post-processed by genesis tracking software (e.g., GFDL vortex tracker), ATCF edeck files containing probability of genesis forecasts, operational shapefile warning areas, and ATCF reference track dataset(s) (e.g., Best Track analysis and CARQ operational tracks). It writes categorical counts and statistics. The capability to modify the spatial and temporal tolerances when matching forecasts to reference genesis events, as well as scoring those matched pairs, gives users the ability to condition the criteria based on model performance and/or conduct sensitivity analyses. Statistical aspects are outlined in :numref:`tc-gen_stat_aspects` and practical aspects of the TC-Gen tool are described in :numref:`tc-gen_practical_info`. .. _tc-gen_stat_aspects: @@ -20,15 +20,15 @@ For deterministic forecasts specified using the **-track** command line option, As with other extreme events (where the event occurs much less frequently than the non-event), the correct negative category is not computed since the non-events would dominate the contingency table. Therefore, only statistics that do not include correct negatives should be considered for this tool. The following CTS statistics are relevant: Base rate (BASER), Mean forecast (FMEAN), Frequency Bias (FBIAS), Probability of Detection (PODY), False Alarm Ratio (FAR), Critical Success Index (CSI), Gilbert Skill Score (GSS), Extreme Dependency Score (EDS), Symmetric Extreme Dependency Score (SEDS), Bias-Adjusted Gilbert Skill Score (BAGSS). -For probabilistic forecasts specified using the **-edeck** command line option, it identifies genesis events in the reference dataset. It applies user-specified configuration options to pair the forecast probabilities to the reference genesis events. These pairs are added to an Nx2 probabilistic contingency table. If the reference genesis event occurs within in the predicted time window, the pair is counted in the observation-yes column. Otherwise, it is added to the observation-no column. +For probabilistic forecasts specified using the **-edeck** command line option, it identifies genesis events in the reference dataset. It applies user-specified configuration options to pair the forecast probabilities to the reference genesis events. These pairs are added to an Nx2 probabilistic contingency table. If the reference genesis event occurs within the predicted time window, the pair is counted in the observation-yes column. Otherwise, it is added to the observation-no column. -For warning area shapefiles specified using the **-shape** command line option, it processes metadata from the corresponding database files. The database file is assumed to exist at exactly the same path as the shapefile, but with a ".dbf" suffix instead of ".shp". Note that only shapefiles exactly following the NOAA National Hurricane Center's (NHC) "gtwo_areas_YYYYMMDDHHMM.shp" file naming and corresonding metadata conventions are supported. For each shapefile record, the database file defines corresponding probability values for one or more time periods. Percentages may be provided for the probability of genesis inside the shape within 2, 5, or 7 days from issuance time that is parsed from the file name. Note that 5 day probabilities were discontinued in 2023. The 2 and 7 day probabilities are provided in database file fields named "PROB2DAY" and "PROB7DAY", respectively. Care is taken to identify and either ignore or update duplicate shapes found in the input. +For warning area shapefiles specified using the **-shape** command line option, it processes metadata from the corresponding database files. The database file is assumed to exist at exactly the same path as the shapefile, but with a ".dbf" suffix instead of ".shp". Note that only shapefiles exactly following the NOAA National Hurricane Center's (NHC) "gtwo_areas_YYYYMMDDHHMM.shp" file naming and corresponding metadata conventions are supported. For each shapefile record, the database file defines corresponding probability values for one or more time periods. Percentages may be provided for the probability of genesis inside the shape within 2, 5, or 7 days from issuance time that is parsed from the file name. Note that 5 day probabilities were discontinued in 2023. The 2 and 7 day probabilities are provided in database file fields named "PROB2DAY" and "PROB7DAY", respectively. Care is taken to identify and either ignore or update duplicate shapes found in the input. -The shapes are then subset based on the filtering criteria in the configuration file. For each probability and shape, the reference genesis events are searched for a match within the defined time window. These pairs are added to an Nx2 probabilistic contingency table. The probabilistic contingeny tables and statistics are computed and reported separately for filter defined and lead hour encountered in the input. +The shapes are then subset based on the filtering criteria in the configuration file. For each probability and shape, the reference genesis events are searched for a match within the defined time window. These pairs are added to an Nx2 probabilistic contingency table. The probabilistic contingency tables and statistics are computed and reported separately for each filter defined and lead hour encountered in the input. Other considerations for interpreting the output of the TC-Gen tool involve the size of the contingency table output. The size of the contingency table will change depending on the number of matches. Additionally, the number of misses is based on the forecast duration and interval (specified in the configuration file). This change is due to the number of model opportunities to forecast the event, which is determined by the specified duration/interval. -Care should be taken when interpreting the statistics for filtered data. In some cases, variables (e.g. storm name) are only available in either the forecast or reference datasets, rather than both. When filtering on a field that is only present in one dataset, the contingency table counts will be impacted. Similarly, the initialization field only impacts the model forecast data. If the valid time (which will impact the reference dataset) isn't also specified, the forecasts will be filtered and matched such that the number of misses will erroneously increase. See :numref:`tc-gen_practical_info` for more detail. +Care should be taken when interpreting the statistics for filtered data. In some cases, variables (e.g., storm name) are only available in either the forecast or reference datasets, rather than both. When filtering on a field that is only present in one dataset, the contingency table counts will be impacted. Similarly, the initialization field only impacts the model forecast data. If the valid time (which will impact the reference dataset) isn't also specified, the forecasts will be filtered and matched such that the number of misses will erroneously increase. See :numref:`tc-gen_practical_info` for more detail. .. _tc-gen_practical_info: @@ -45,10 +45,10 @@ The usage statement for tc_gen is shown below: .. code-block:: none Usage: tc_gen - -genesis source - -edeck source - -shape source - -track source + -genesis path + -edeck path + -shape path + -track path -config file [-out base] [-log file] @@ -59,15 +59,15 @@ TC-Gen has three required arguments and accepts optional ones. Required Arguments for tc_gen ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ -1. The **-genesis source** argument is the path to one or more ATCF or fort.66 (see documentation listed below) files generated by the Geophysical Fluid Dynamics Laboratory (GFDL) Vortex Tracker when run in tcgen mode or an ASCII file list or a top-level directory containing them. The required file format is described in the "Output formats" section of the `GFDL Vortex Tracker users guide. `_ +1. The **-genesis path** argument is the path to one or more ATCF or fort.66 (see documentation listed below) files generated by the Geophysical Fluid Dynamics Laboratory (GFDL) Vortex Tracker when run in tcgen mode or an ASCII file list or a top-level directory containing them. The required file format is described in the "Output formats" section of the `GFDL Vortex Tracker users guide. `_ -2. The **-edeck source** argument is the path to one or more ATCF edeck files, an ASCII file list containing them, or a top-level directory with files matching the regular expression ".dat". The probability of genesis are read from each edeck input file and verified against at the **-track** data. +2. The **-edeck path** argument is the path to one or more ATCF edeck files, an ASCII file list containing them, or a top-level directory with files matching the regular expression ".dat". The probability of genesis are read from each edeck input file and verified against the **-track** data. -3. The **-shape source** argument is the path to one or more NHC genesis warning area shapefiles, an ASCII file list containing them, or a top-level directory with files matching the regular expression "gtwo_areas.*.shp". The genesis warning areas and corresponding forecast probability values area verified against the **-track** data. +3. The **-shape path** argument is the path to one or more NHC genesis warning area shapefiles, an ASCII file list containing them, or a top-level directory with files matching the regular expression "gtwo_areas.*.shp". The genesis warning areas and corresponding forecast probability values are verified against the **-track** data. -Note: At least one of the **-genesis**, **-edeck**, or **-shape** command line options are required. +Note: At least one of the **-genesis**, **-edeck**, or **-shape** command line options is required. -4. The **-track source** argument is one or more ATCF reference track files or an ASCII file list or top-level directory containing them, with files ending in ".dat". This tool processes either Best track data from bdeck files, or operational track data (e.g. CARQ) from adeck files, or both. Providing both bdeck and adeck files will result in a richer dataset to match with the **-genesis** files. Both adeck and bdeck data should be provided using the **-track** option. The **-track** option must be used at least once. +4. The **-track path** argument is one or more ATCF reference track files or an ASCII file list or top-level directory containing them, with files ending in ".dat". This tool processes either Best track data from bdeck files, or operational track data (e.g., CARQ) from adeck files, or both. Providing both bdeck and adeck files will result in a richer dataset to match with the **-genesis** files. Both adeck and bdeck data should be provided using the **-track** option. The **-track** option must be used at least once. 5. The **-config** file argument indicates the name of the configuration file to be used. The contents of the configuration file are discussed below. @@ -89,73 +89,73 @@ The TC-Gen tool implements the following logic: * For **-track** inputs: - * Parse the forecast genesis data and identify forecast genesis events separately for each model present. + * Parse the forecast genesis data and identify forecast genesis events separately for each model present. - * Loop over the filters defined in the configuration file and apply the following logic for each. + * Loop over the filters defined in the configuration file and apply the following logic for each. - * For each Best track genesis event meeting the filter critera, determine the initialization and lead times for which the model had an opportunity to forecast that genesis event. Store an unmatched genesis pair for each case. + * For each Best track genesis event meeting the filter criteria, determine the initialization and lead times for which the model had an opportunity to forecast that genesis event. Store an unmatched genesis pair for each case. - * For each forecast genesis event, search for a matching Best track. A configurable boolean option controls whether all Best track points are considered for a match or only the single Best track genesis point. A match occurs if the Best track point valid time is within a configurable window around the forecast genesis time and the Best track point location is within a configurable radius of the forecast genesis location. If a Best track match is found, store the storm ID. + * For each forecast genesis event, search for a matching Best track. A configurable boolean option controls whether all Best track points are considered for a match or only the single Best track genesis point. A match occurs if the Best track point valid time is within a configurable window around the forecast genesis time and the Best track point location is within a configurable radius of the forecast genesis location. If a Best track match is found, store the storm ID. - * If no Best track match is found, apply the same logic to search the operational track points with lead time of 0 hours. If an operational match is found, store the storm ID. + * If no Best track match is found, apply the same logic to search the operational track points with lead time of 0 hours. If an operational match is found, store the storm ID. - * If a matching storm ID is found, match the forecast genesis event to the Best track genesis event for that storm ID. + * If a matching storm ID is found, match the forecast genesis event to the Best track genesis event for that storm ID. - * If no matching storm ID is found, store an unmatched pair for the genesis forecast. + * If no matching storm ID is found, store an unmatched pair for the genesis forecast. - * Loop through the genesis pairs and populate contingency tables using two methods, the development (dev) and operational (ops) methods. For each pair, if the forecast genesis event is unmatched, score it as a dev and ops FALSE ALARM. If the Best track genesis event is unmatched, score it as a dev and ops MISS. Score each matched genesis pair as follows: + * Loop through the genesis pairs and populate contingency tables using two methods, the development (dev) and operational (ops) methods. For each pair, if the forecast genesis event is unmatched, score it as a dev and ops FALSE ALARM. If the Best track genesis event is unmatched, score it as a dev and ops MISS. Score each matched genesis pair as follows: - * If the forecast initialization time is at or after the Best track genesis event, DISCARD this case and exclude it from the statistics. + * If the forecast initialization time is at or after the Best track genesis event, DISCARD this case and exclude it from the statistics. - * Compute the difference between the forecast and Best track genesis events in time and space. If they are both within the configurable tolerance, score it as a dev HIT. If not, score it as a dev FALSE ALARM. + * Compute the difference between the forecast and Best track genesis events in time and space. If they are both within the configurable tolerance, score it as a dev HIT. If not, score it as a dev FALSE ALARM. - * Compute the difference between the Best track genesis time and model initialization time. If it is within the configurable tolerance, score it as an ops HIT. If not, score it as an ops FALSE ALARM. + * Compute the difference between the Best track genesis time and model initialization time. If it is within the configurable tolerance, score it as an ops HIT. If not, score it as an ops FALSE ALARM. - * Do not count any CORRECT NEGATIVES. + * Do not count any CORRECT NEGATIVES. - * Report the contingency table hits, misses, and false alarms separately for each forecast model and configuration file filter. The development (dev) scoring method is indicated in the output as *GENESIS_DEV* while the operational (ops) scoring method is indicated as *GENESIS_OPS*. + * Report the contingency table hits, misses, and false alarms separately for each forecast model and configuration file filter. The development (dev) scoring method is indicated in the output as *GENESIS_DEV* while the operational (ops) scoring method is indicated as *GENESIS_OPS*. * For **-edeck** inputs: - * Parse the ATCF edeck files. Ignore any lines not containing "GN" and "genFcst", which indicate a genesis probability forecast. Also, ignore any lines which do not contain a predicted genesis location (latitude and longitude) or genesis time. + * Parse the ATCF edeck files. Ignore any lines not containing "GN" and "genFcst", which indicate a genesis probability forecast. Also, ignore any lines which do not contain a predicted genesis location (latitude and longitude) or genesis time. - * Loop over the filters defined in the configuration file and apply the following logic for each. + * Loop over the filters defined in the configuration file and apply the following logic for each. - * Subset the genesis probability forecasts based on the current filter criteria. Typically, genesis probability forecast are provided for multiple lead times. Create separate Nx2 probabilistic contingency tables for each unique combination of predicted lead time and model name. + * Subset the genesis probability forecasts based on the current filter criteria. Typically, genesis probability forecasts are provided for multiple lead times. Create separate Nx2 probabilistic contingency tables for each unique combination of predicted lead time and model name. - * For each genesis probability forecast, search for a matching Best track. A configurable boolean option controls whether all Best track points are considered for a match or only the single Best track genesis point. A match occurs if the Best track point valid time is within a configurable window around the forecast genesis time and the Best track point location is within a configurable radius of the forecast genesis location. If a Best track match is found, store the storm ID. + * For each genesis probability forecast, search for a matching Best track. A configurable boolean option controls whether all Best track points are considered for a match or only the single Best track genesis point. A match occurs if the Best track point valid time is within a configurable window around the forecast genesis time and the Best track point location is within a configurable radius of the forecast genesis location. If a Best track match is found, store the storm ID. - * If no Best track match is found, apply the same logic to search the operational track points with lead time of 0 hours. If an operational match is found, store the storm ID. + * If no Best track match is found, apply the same logic to search the operational track points with lead time of 0 hours. If an operational match is found, store the storm ID. - * If no matching storm ID is found, add the unmatched forecast to the observation-no column of the Nx2 probabilistic contingency table. + * If no matching storm ID is found, add the unmatched forecast to the observation-no column of the Nx2 probabilistic contingency table. - * If a matching storm ID is found, check whether that storm's genesis occurred within the predicted time window: between the forecast initialization time and the predicted lead time. If so, add the matched forecast to the observation-yes column. If not, add it to observation-no column. + * If a matching storm ID is found, check whether that storm's genesis occurred within the predicted time window: between the forecast initialization time and the predicted lead time. If so, add the matched forecast to the observation-yes column. If not, add it to the observation-no column. - * Report the Nx2 probabilistic contingency table counts and statistics for each forecast model, lead time, and configuration file filter. These counts and statistics are identified in the output files as *PROB_GENESIS*. + * Report the Nx2 probabilistic contingency table counts and statistics for each forecast model, lead time, and configuration file filter. These counts and statistics are identified in the output files as *PROB_GENESIS*. * For **-shape** inputs: - * For each input shapefile, parse the timestamp from the "gtwo_areas_YYYYMMDDHHMM.shp" naming convention, and error out otherwise. Round the timestamp to the nearest synoptic time (e.g. 00, 06, 12, 18) and store that as the issuance time. + * For each input shapefile, parse the timestamp from the "gtwo_areas_YYYYMMDDHHMM.shp" naming convention, and error out otherwise. Round the timestamp to the nearest synoptic time (e.g., 00, 06, 12, 18) and store that as the issuance time. - * Open the shapefile and corresponding database file. Process each record. + * Open the shapefile and corresponding database file. Process each record. - * For each record, extract the shape and metadata which defines the basin and 2, 5, and 7 day probabilities. + * For each record, extract the shape and metadata which defines the basin and 2, 5, and 7 day probabilities. - * Check if this shape is a duplicate that has already been processed. If it is an exact duplicate, with the same basin, file timestamp, issue time, and min/max lat/lon values, ignore it. If the file timestamp is older than the existing shape, also ignore it. If the file timestamp is newer than the existing shape, replace the existing shape with the new one. + * Check if this shape is a duplicate that has already been processed. If it is an exact duplicate, with the same basin, file timestamp, issue time, and min/max lat/lon values, ignore it. If the file timestamp is older than the existing shape, also ignore it. If the file timestamp is newer than the existing shape, replace the existing shape with the new one. - * Loop over the filters defined in the configuration file and apply the following logic for each. + * Loop over the filters defined in the configuration file and apply the following logic for each. - * Subset the list of genesis shapes based on the current filter criteria. + * Subset the list of genesis shapes based on the current filter criteria. - * Search the Best track genesis events to see if any occurred inside the shape within 7 days of the issuance time. If multiple genesis events occurred, choose the one closest to the issuance time. + * Search the Best track genesis events to see if any occurred inside the shape within 7 days of the issuance time. If multiple genesis events occurred, choose the one closest to the issuance time. - * If not found, score each probability as a miss. + * If not found, score each probability as a miss. - * If found, further check the 2 and 5 day time windows to classify each probability as a hit or miss. + * If found, further check the 2 and 5 day time windows to classify each probability as a hit or miss. - * Add each probability pair to an Nx2 probabilistic contingency table, tracking results separately for each lead time. + * Add each probability pair to an Nx2 probabilistic contingency table, tracking results separately for each lead time. - * Report the Nx2 probabilistic contingency table counts and statistics for each lead time. These counts and statistics are identified in the output files as *GENESIS_SHAPE*. + * Report the Nx2 probabilistic contingency table counts and statistics for each lead time. These counts and statistics are identified in the output files as *GENESIS_SHAPE*. tc_gen Configuration File ------------------------- @@ -313,7 +313,7 @@ ______________________ basin_mask = []; -The **basin_mask** entry is an array of strings listing tropical cycline basin abbreviations (e.g. AL, EP, CP, WP, NI, SI, AU, and SP). The configuration entry **basin_file** defines the path to a NetCDF file which defines these regions. The default file (**basin_global_tenth_degree.nc**) is bundled with MET. If **basin_mask** is left empty, genesis events for all basins will be included. If non-empty, the union of specified basins will be used. If **vx_mask** is also specified, the analysis is done on the intersection of those masking areas. +The **basin_mask** entry is an array of strings listing tropical cyclone basin abbreviations (e.g., AL, EP, CP, WP, NI, SI, AU, and SP). The configuration entry **basin_file** defines the path to a NetCDF file which defines these regions. The default file (**basin_global_tenth_degree.nc**) is bundled with MET. If **basin_mask** is left empty, genesis events for all basins will be included. If non-empty, the union of specified basins will be used. If **vx_mask** is also specified, the analysis is done on the intersection of those masking areas. The **vx_mask** and **basin_mask** names are concatenated and written to the **VX_MASK** output column. @@ -333,7 +333,7 @@ ______________________ genesis_match_point_to_track = TRUE; -The **genesis_match_point_to_track** entry is a boolean which controls the matching logic. When set to its default value of TRUE, for each forecast genesis event, all Best track points are searched for a match. This logic implements the method used by the NOAA National Hurricane Center. When set to FALSE, only the single Best track genesis point is considered for a match. When selecting FALSE, users are encouraged to adjust the **genesis_match_radius** and/or **gensesis_match_window** options, described below, to enable matches to be found. +The **genesis_match_point_to_track** entry is a boolean which controls the matching logic. When set to its default value of TRUE, for each forecast genesis event, all Best track points are searched for a match. This logic implements the method used by the NOAA National Hurricane Center. When set to FALSE, only the single Best track genesis point is considered for a match. When selecting FALSE, users are encouraged to adjust the **genesis_match_radius** and/or **genesis_match_window** options, described below, to enable matches to be found. ______________________ @@ -352,7 +352,7 @@ ______________________ end = 0; } -The **genesis_match_window** entry defines a time window, in hours, relative to the forecast genesis time. When searching for a match, only Best or operational tracks with a track point falling within this time window will be considered. The default time window of 0 requires a Best or operational track to exist at the forecast genesis time for a match to be found. Increasing this time window should lead to an increase in the number matched genesis pairs. For example, setting *end = 12;* would allow forecast genesis events to match Best tracks up to 12 hours prior to their existence. +The **genesis_match_window** entry defines a time window, in hours, relative to the forecast genesis time. When searching for a match, only Best or operational tracks with a track point falling within this time window will be considered. The default time window of 0 requires a Best or operational track to exist at the forecast genesis time for a match to be found. Increasing this time window should lead to an increase in the number of matched genesis pairs. For example, setting *end = 12;* would allow forecast genesis events to match Best tracks up to 12 hours prior to their existence. ______________________ @@ -382,7 +382,7 @@ ______________________ end = 48; } -The **ops_hit_window** entry defines a time window, in hours, relative to the Best track genesis time. The model initialization time for the forecast genesis event must occur within this time window for the pairs to be counted as a contingency table HIT for the operationl scoring method. Otherwise, the pair is counted as a FALSE ALARM. +The **ops_hit_window** entry defines a time window, in hours, relative to the Best track genesis time. The model initialization time for the forecast genesis event must occur within this time window for the pairs to be counted as a contingency table HIT for the operational scoring method. Otherwise, the pair is counted as a FALSE ALARM. ______________________ @@ -390,7 +390,7 @@ ______________________ discard_init_post_genesis_flag = TRUE; -The **discard_init_post_genesis_flag** entry is a boolean which indicates whether or not forecast genesis events from model intializations occurring at or after the matching Best track genesis time should be discarded. If true, those cases are not scored in the contingency table. If false, they are included in the counts. +The **discard_init_post_genesis_flag** entry is a boolean which indicates whether or not forecast genesis events from model initializations occurring at or after the matching Best track genesis time should be discarded. If true, those cases are not scored in the contingency table. If false, they are included in the counts. ______________________ @@ -417,7 +417,7 @@ ______________________ best_fn_oy = TRUE; } -The **nc_pairs_flag** entry is a dictionary of booleans indicating which fields should be written to the NetCDF genesis pairs output file. Each type of output is enabled by setting it to TRUE and disabled by setting it to FALSE. The **latlon** option writes the latitude and longitude values of the output grid. The remaining options write a count of the number of points occuring within each grid cell. The **fcst_genesis** and **best_genesis** options write counts of the forecast and Best track genesis locations. The **fcst_track** and **best_track** options write counts of the full set of track point locations, which can be refined by the **valid_minus_genesis_diff_thresh** option, described below. The **fcst_fy_oy** and **fcst_fy_on** options write counts for the locations of forecast genesis event HITS and FALSE ALARMS. The **best_fy_oy** and **best_fn_oy** options write counts for the locations of Best track genesis event HITS and MISSES. Note that since matching forecast and Best track genesis events may occur in different grid cells, their counts are reported separately. +The **nc_pairs_flag** entry is a dictionary of booleans indicating which fields should be written to the NetCDF genesis pairs output file. Each type of output is enabled by setting it to TRUE and disabled by setting it to FALSE. The **latlon** option writes the latitude and longitude values of the output grid. The remaining options write a count of the number of points occurring within each grid cell. The **fcst_genesis** and **best_genesis** options write counts of the forecast and Best track genesis locations. The **fcst_tracks** and **best_tracks** options write counts of the full set of track point locations, which can be refined by the **valid_minus_genesis_diff_thresh** option, described below. The **fcst_fy_oy** and **fcst_fy_on** options write counts for the locations of forecast genesis event HITS and FALSE ALARMS. The **best_fy_oy** and **best_fn_oy** options write counts for the locations of Best track genesis event HITS and MISSES. Note that since matching forecast and Best track genesis events may occur in different grid cells, their counts are reported separately. ______________________ @@ -426,7 +426,7 @@ ______________________ valid_minus_genesis_diff_thresh = NA; -The **valid_minus_genesis_diff_thresh** is a threshold which affects the counts in the NetCDF pairs output file. The fcst_tracks and best_tracks options, described above, turn on counts for the forecast and Best track points. This option defines which of those track points should be counted by thresholding the track point valid time minus genesis time difference. If set to NA, the default threshold which always evaluates to true, all track points will be counted. Setting <=0 would count the genesis point and all track points prior. Setting >0 would count all points after genesis. And setting >=-12||<=12 would could all points within 12 hours of the genesis time. +The **valid_minus_genesis_diff_thresh** is a threshold which affects the counts in the NetCDF pairs output file. The fcst_tracks and best_tracks options, described above, turn on counts for the forecast and Best track points. This option defines which of those track points should be counted by thresholding the track point valid time minus genesis time difference. If set to NA, the default threshold which always evaluates to true, all track points will be counted. Setting <=0 would count the genesis point and all track points prior. Setting >0 would count all points after genesis. And setting >=-12||<=12 would count all points within 12 hours of the genesis time. ______________________ diff --git a/docs/Users_Guide/tc-pairs.rst b/docs/Users_Guide/tc-pairs.rst index 38ecbb8c76..d1515db6a1 100644 --- a/docs/Users_Guide/tc-pairs.rst +++ b/docs/Users_Guide/tc-pairs.rst @@ -27,9 +27,9 @@ TC diagnostics provide information about a TC's structure or its environment. Ea * Satellite-based diagnostics provide information about the storm structure as observed by geostationary satellite infrared imagery. Examples include information about the shape and extent of the cold-cirrus canopy of the TC and whether patterns are present that may portend intensification. -Diagnostics are critically important for training and running statistical-dynamical models that predict a TC's intensity or size. One of the most well-known diagnostics sets is that of the Statistical Hurricane Intensity Prediction Scheme (SHIPS), which supports predictions of TC intensity. A large 30-year development dataset of TC diagnostics has been retrospectively derived to support the training of the SHIPS intensity model as well as other related models such as the Logistic Growth Equation Model (LGEM), SHIPS Rapid Intensification Index (SHIPS-RII), and others. These diagnostics, called *lsdiag* for "large scale" environment, are computed using a *perfect prog* approach in which the diagnostics are computed on the reference model's verifying analyses to generate a set of time-dependent diagnostics from t=0 out to the desired maximum forecast lead time. This is repeated for each initialization, building up a full history of diagnostics for each storm. By using the subsequent verifying analysis for later lead times, the model is taken to be "perfect", removing the impact of model forecast errors. The resulting developmental dataset is ideal for training statistical-dynamical models such as SHIPS. To generate forecasts in real-time, the diagnostics are computed along a forecast track (often taken to be the National Hurricane Center's official forecast) using the fields of the underlying NWP model (e.g, the Global Forecast System, or GFS model). The resulting diagnostics are then used as *predictors* in models like SHIPS and LGEM to predict a TC's future intensity or probability of undergoing rapid intensification. +Diagnostics are critically important for training and running statistical-dynamical models that predict a TC's intensity or size. One of the most well-known diagnostics sets is that of the Statistical Hurricane Intensity Prediction Scheme (SHIPS), which supports predictions of TC intensity. A large 30-year development dataset of TC diagnostics has been retrospectively derived to support the training of the SHIPS intensity model as well as other related models such as the Logistic Growth Equation Model (LGEM), SHIPS Rapid Intensification Index (SHIPS-RII), and others. These diagnostics, called *lsdiag* for "large scale" environment, are computed using a *perfect prog* approach in which the diagnostics are computed on the reference model's verifying analyses to generate a set of time-dependent diagnostics from t=0 out to the desired maximum forecast lead time. This is repeated for each initialization, building up a full history of diagnostics for each storm. By using the subsequent verifying analysis for later lead times, the model is taken to be "perfect", removing the impact of model forecast errors. The resulting developmental dataset is ideal for training statistical-dynamical models such as SHIPS. To generate forecasts in real-time, the diagnostics are computed along a forecast track (often taken to be the National Hurricane Center's official forecast) using the fields of the underlying NWP model (e.g., the Global Forecast System, or GFS model). The resulting diagnostics are then used as *predictors* in models like SHIPS and LGEM to predict a TC's future intensity or probability of undergoing rapid intensification. -Beside their use in TC prediction, TC diagnostics can be very useful to forecasters to understand the forecast scenario. They are also useful to model developers for evaluation of model errors and understanding model performance under different environmental conditions. For instance, a modeler may wish to understand their model's track biases under conditions of high vertical wind shear. TC diagnostics can also be used to understand the sensitivity of the model's intensity predictions to oceanic conditions such as upwelling. The TC-Pairs tool allows filtering and subsetting based on the values of one or several TC diagnostics. +Besides their use in TC prediction, TC diagnostics can be very useful to forecasters to understand the forecast scenario. They are also useful to model developers for evaluation of model errors and understanding model performance under different environmental conditions. For instance, a modeler may wish to understand their model's track biases under conditions of high vertical wind shear. TC diagnostics can also be used to understand the sensitivity of the model's intensity predictions to oceanic conditions such as upwelling. The TC-Pairs tool allows filtering and subsetting based on the values of one or several TC diagnostics. As of MET v11.0.0, two types of TC diagnostics are supported in TC-Pairs: @@ -89,7 +89,7 @@ Optional Arguments for tc_pairs 5. The **-diag source path** argument indicates the TC-Pairs acceptable format data containing the tropical cyclone diagnostics dataset corresponding to the adeck tracks. The **source** can be set to CIRA_DIAG_RT or SHIPS_DIAG_RT to indicate the input diagnostics data source. The **path** argument specifies the name of a TC-Pairs acceptable format file or top-level directory containing TC-Pairs acceptable format files ending in ".dat" to be processed. Support for additional diagnostic sources will be added in future releases. -6. The -**out base** argument indicates the path of the output file base. This argument overrides the default output file base (**./out_tcmpr**). +6. The **-out base** argument indicates the path of the output file base. This argument overrides the default output file base (**./out_tcmpr**). 7. The **-log file** option directs output and errors to the specified log file. All messages will be written to that file as well as standard out and error. Thus, users can save the messages without having to redirect the output on the command line. The default behavior is no log file. @@ -179,7 +179,7 @@ ____________________ check_dup = FALSE; -The **check_dup** flag expects either TRUE and FALSE, indicating whether the code should check for duplicate ATCF lines when building tracks. Setting **check_dup** to TRUE will check for duplicated lines, and produce output information regarding the duplicate. Any duplicated ATCF line will not be processed in the tc_pairs output. Setting **check_dup** to FALSE, will still exclude tracks that decrease with time, and will overwrite repeated lines, but specific duplicate log information will not be output. Setting **check_dup** to FALSE will make parsing the track quicker. +The **check_dup** flag expects either TRUE or FALSE, indicating whether the code should check for duplicate ATCF lines when building tracks. Setting **check_dup** to TRUE will check for duplicated lines, and produce output information regarding the duplicate. Any duplicated ATCF line will not be processed in the tc_pairs output. Setting **check_dup** to FALSE, will still exclude tracks that decrease with time, and will overwrite repeated lines, but specific duplicate log information will not be output. Setting **check_dup** to FALSE will make parsing the track quicker. ____________________ @@ -187,7 +187,7 @@ ____________________ interp12 = NONE; -The **interp12** flag expects the entry NONE, FILL, or REPLACE, indicating whether special processing should be performed for interpolated forecasts. The NONE option indicates no changes are made to the interpolated forecasts. The FILL and REPLACE (default) options determine when the 12-hour interpolated forecast (normally indicated with a "2" or "3" at the end of the ATCF ID) will be renamed with the 6-hour interpolated ATCF ID (normally indicated with the letter "I" at the end of the ATCF ID). The FILL option renames the 12-hour interpolated forecasts with the 6-hour interpolated forecast ATCF ID only when the 6-hour interpolated forecasts is missing (in the case of a 6-hour interpolated forecast which only occurs every 12-hours (e.g. EMXI, EGRI), the 6-hour interpolated forecasts will be "filled in" with the 12-hour interpolated forecasts in order to provide a record every 6-hours). The REPLACE option renames all 12-hour interpolated forecasts with the 6-hour interpolated forecasts ATCF ID regardless of whether the 6-hour interpolated forecast exists. The original 12-hour ATCF ID will also be retained in the output file (all modified ATCF entries will appear at the end of the TC-Pairs output file). This functionality expects both the 12-hour and 6-hour early (interpolated) ATCF IDs to be listed in the model field. +The **interp12** flag expects the entry NONE, FILL, or REPLACE, indicating whether special processing should be performed for interpolated forecasts. The NONE option indicates no changes are made to the interpolated forecasts. The FILL and REPLACE (default) options determine when the 12-hour interpolated forecast (normally indicated with a "2" or "3" at the end of the ATCF ID) will be renamed with the 6-hour interpolated ATCF ID (normally indicated with the letter "I" at the end of the ATCF ID). The FILL option renames the 12-hour interpolated forecasts with the 6-hour interpolated forecast ATCF ID only when the 6-hour interpolated forecast is missing (in the case of a 6-hour interpolated forecast which only occurs every 12-hours (e.g., EMXI, EGRI), the 6-hour interpolated forecasts will be "filled in" with the 12-hour interpolated forecasts in order to provide a record every 6-hours). The REPLACE option renames all 12-hour interpolated forecasts with the 6-hour interpolated forecasts ATCF ID regardless of whether the 6-hour interpolated forecast exists. The original 12-hour ATCF ID will also be retained in the output file (all modified ATCF entries will appear at the end of the TC-Pairs output file). This functionality expects both the 12-hour and 6-hour early (interpolated) ATCF IDs to be listed in the model field. ____________________ @@ -208,7 +208,7 @@ ____________________ The **consensus** array allows users to derive consensus forecasts from any number of models. A consensus forecast is computed as the average intensity and location of the members which comprise it. TC-Pairs attempts to derive consensus forecasts for each unique storm ID and initialization time found in the input track data. Each array entry is a dictionary which defines the consensus name, membership, and requirements: - The **name** field is a string defining the consensus model name to be written. -- The **members** field is a comma-separated array of model ID stings which define the members of the consensus. +- The **members** field is a comma-separated array of model ID strings which define the members of the consensus. - The **required** field is a comma-separated array of true/false values associated with each consensus member. If a member is designated as true, that member must be present in order for the consensus to be generated. If a member is false, the consensus will be generated regardless of whether or not the member is present. The required array can either be empty or have the same length as the members array. If empty, it defaults to all false. - The **min_req** field is the number of members required in order for the consensus to be computed. The **required** and **min_req** field options are applied at each forecast lead time. If any member of the consensus has a non-valid position or intensity value, the consensus for that valid time will not be generated. - Tropical cyclone diagnostics, if provided on the command line, are included in the computation of consensus tracks. The consensus diagnostics are computed as the mean of the diagnostics for the members. The **diag_required** and **min_diag_req** entries apply the same logic described above, but to the computation of each consensus diagnostic value rather than the consensus track location and intensity. If **diag_required** is missing or an empty list, it defaults to all false. If **min_diag_req** is missing, it defaults to 0. @@ -286,10 +286,10 @@ ____________________ .. code-block:: none - watch_warn = { - file_name = "MET_BASE/tc_data/wwpts_us.txt"; - time_offset = -14400; - } + watch_warn = { + file_name = "MET_BASE/tc_data/wwpts_us.txt"; + time_offset = -14400; + } The **watch_warn** field specifies the file name and time applied offset to the **watch_warn** flag. The **file_name** string specifies the path of the watch/warning file to be used to determine when a watch or warning is in effect during the forecast initialization and verification times. The default file is named **wwpts_us.txt**, which is found in the installed *share/met/tc_data/* directory within the MET build. The **time_offset** string is the time window (in seconds) assigned to the watch/warning. Due to the non-uniform time watches and warnings are issued, a time window is assigned for which watch/warnings are included in the verification for each valid time. The default watch/warn file is static, and therefore may not include warned storms beyond the current MET code release date; therefore users may wish to create a post in the `METplus GitHub Discussions Forum `_ in order to obtain the most recent watch/warning file if the static file does not contain storms of interest. @@ -297,24 +297,24 @@ ____________________ .. code-block:: none - diag_info_map = [ - { - diag_source = "CIRA_DIAG_RT"; - track_source = "GFS"; - field_source = "GFS_0p50"; - match_to_track = []; - diag_name = []; - }, - { - diag_source = "SHIPS_DIAG_RT"; - track_source = "SHIPS_TRK"; - field_source = "GFS_0p50"; - match_to_track = [ "OFCL" ]; - diag_name = []; - } - ]; - -A TCMPR line is written to the output for each track point. If diagnostics data is also defined for that track point, a TCDIAG line is written immediately after the corresponding TCMPR line. The contents of that TCDIAG line is determined by the **diag_info_map** entry. + diag_info_map = [ + { + diag_source = "CIRA_DIAG_RT"; + track_source = "GFS"; + field_source = "GFS_0p50"; + match_to_track = []; + diag_name = []; + }, + { + diag_source = "SHIPS_DIAG_RT"; + track_source = "SHIPS_TRK"; + field_source = "GFS_0p50"; + match_to_track = [ "OFCL" ]; + diag_name = []; + } + ]; + +A TCMPR line is written to the output for each track point. If diagnostics data is also defined for that track point, a TCDIAG line is written immediately after the corresponding TCMPR line. The contents of that TCDIAG line are determined by the **diag_info_map** entry. The **diag_info_map** entries define how the diagnostics read with the **-diag** command line option should be used. Each array element is a dictionary consisting of entries for **diag_source**, **track_source**, **field_source**, **match_to_track**, and **diag_name**. @@ -328,44 +328,44 @@ ____________________ .. code-block:: none - diag_convert_map = [ - { - diag_source = "CIRA_DIAG"; - key = [ "(10C)", "(10KT)", "(10M/S)" ]; - convert(x) = x / 10; - }, - { - diag_source = "SHIPS_DIAG"; - key = [ "LAT", "LON", "CSST", "RSST", "DSST", "DSTA", "XDST", "XNST", "NSST", "NSTA", - "NTMX", "NTFR", "U200", "U20C", "V20C", "E000", "EPOS", "ENEG", "EPSS", "ENSS", - "T000", "TLAT", "TLON", "TWAC", "TWXC", "G150", "G200", "G250", "V000", "V850", - "V500", "V300", "SHDC", "SHGC", "T150", "T200", "T250", "SHRD", "SHRS", "SHRG", - "HE07", "HE05", "PW01", "PW02", "PW03", "PW04", "PW05", "PW06", "PW07", "PW08", - "PW09", "PW10", "PW11", "PW12", "PW13", "PW14", "PW15", "PW16", "PW17", "PW18", - "PW20", "PW21" ]; - convert(x) = x / 10; - }, - { - diag_source = "SHIPS_DIAG"; - key = [ "VVAV", "VMFX", "VVAC" ]; - convert(x) = x / 100; - }, - { + diag_convert_map = [ + { + diag_source = "CIRA_DIAG"; + key = [ "(10C)", "(10KT)", "(10M/S)" ]; + convert(x) = x / 10; + }, + { + diag_source = "SHIPS_DIAG"; + key = [ "LAT", "LON", "CSST", "RSST", "DSST", "DSTA", "XDST", "XNST", "NSST", "NSTA", + "NTMX", "NTFR", "U200", "U20C", "V20C", "E000", "EPOS", "ENEG", "EPSS", "ENSS", + "T000", "TLAT", "TLON", "TWAC", "TWXC", "G150", "G200", "G250", "V000", "V850", + "V500", "V300", "SHDC", "SHGC", "T150", "T200", "T250", "SHRD", "SHRS", "SHRG", + "HE07", "HE05", "PW01", "PW02", "PW03", "PW04", "PW05", "PW06", "PW07", "PW08", + "PW09", "PW10", "PW11", "PW12", "PW13", "PW14", "PW15", "PW16", "PW17", "PW18", + "PW20", "PW21" ]; + convert(x) = x / 10; + }, + { + diag_source = "SHIPS_DIAG"; + key = [ "VVAV", "VMFX", "VVAC" ]; + convert(x) = x / 100; + }, + { + diag_source = "SHIPS_DIAG"; + key = [ "TADV" ]; + convert(x) = x / 1000000; + }, + { diag_source = "SHIPS_DIAG"; - key = [ "TADV" ]; - convert(x) = x / 1000000; - }, - { - diag_source = "SHIPS_DIAG"; - key = [ "Z850", "D200", "TGRD", "DIVC" ]; - convert(x) = x / 10000000; - }, - { - diag_source = "SHIPS_DIAG"; - key = [ "PENC", "PENV" ]; - convert(x) = x / 10 + 1000; - } - ]; + key = [ "Z850", "D200", "TGRD", "DIVC" ]; + convert(x) = x / 10000000; + }, + { + diag_source = "SHIPS_DIAG"; + key = [ "PENC", "PENV" ]; + convert(x) = x / 10 + 1000; + } + ]; The **diag_convert_map** entries define conversion functions to be applied to diagnostics data read with the **-diag** command line option. Each array element is a dictionary consisting of a **diag_source**, **key**, and **convert(x)** entry. @@ -731,7 +731,7 @@ TC-Pairs produces output in TCST format. The default output file name can be ove - Integer * - 21 - DIAG_i - - Name of the of the ith storm diagnostic (repeated) + - Name of the ith storm diagnostic (repeated) - String * - 22 - VALUE_i diff --git a/docs/Users_Guide/tc-rmw.rst b/docs/Users_Guide/tc-rmw.rst index 8fa47fd98a..ea8415d50d 100644 --- a/docs/Users_Guide/tc-rmw.rst +++ b/docs/Users_Guide/tc-rmw.rst @@ -7,7 +7,7 @@ TC-RMW Tool Introduction ============ -The TC-RMW tool regrids tropical cyclone model data onto a moving range-azimuth grid centered on points along the storm track provided in ATCF format. It can process forecast storm tracks found in ATCF adeck files or analysis tracks (e.g. BEST track) found in ATCF bdeck files. The radial grid spacing can be defined in kilometers or as a factor of the radius of maximum winds (RMW). The azimuthal grid spacing is defined in degrees clockwise from due east. If wind vector fields are specified in the configuration file, the radial and tangential wind components will be computed. Any regridding method available in MET can be used to interpolate data on the model output grid to the specified range-azimuth grid. The regridding will be done separately on each vertical level. +The TC-RMW tool regrids tropical cyclone model data onto a moving range-azimuth grid centered on points along the storm track provided in ATCF format. It can process forecast storm tracks found in ATCF adeck files or analysis tracks (e.g., BEST track) found in ATCF bdeck files. The radial grid spacing can be defined in kilometers or as a factor of the radius of maximum winds (RMW). The azimuthal grid spacing is defined in degrees clockwise from due east. If wind vector fields are specified in the configuration file, the radial and tangential wind components will be computed. Any regridding method available in MET can be used to interpolate data on the model output grid to the specified range-azimuth grid. The regridding will be done separately on each vertical level. Each run of TC-RMW processes a single track. Users should define track filtering criteria in the TC-RMW configuration file to select a single one. While earlier versions of TC-RMW required that the gridded input data files and filtered track points must exactly coincide, that requirement has been relaxed. However output is only written for track points for which gridded input data is provided. @@ -23,7 +23,7 @@ The following sections describe the usage statement, required arguments, and opt Usage: tc_rmw -data file_1 ... file_n | file_list - -deck file + -deck path -config file -out file [-log file] @@ -36,7 +36,7 @@ Required Arguments for tc_rmw 1. The **-data file_1 ... file_n | file_list** option specifies the gridded data files or an ASCII file containing a list of files to be used, as described in :numref:`ascii_file_lists`. -2. The **-deck source** argument is the ATCF format data source. +2. The **-deck path** argument is the ATCF format data source. 3. The **-config file** argument is the configuration file to be used. The contents of the configuration file are discussed below. @@ -179,17 +179,17 @@ _______________________ tangential_velocity_long_field_name = "Tangential Velocity"; -The **tangential_velocity_field_name** and **tangential_velocity_long_field_name** parameters define the field names to give the output tangential velocity grid in the netCDF output file. The parameters are used only if **compute_tangential_and_radial_winds** is set to TRUE. +The **tangential_velocity_field_name** and **tangential_velocity_long_field_name** parameters define the field names to give the output tangential velocity grid in the NetCDF output file. The parameters are used only if **compute_tangential_and_radial_winds** is set to TRUE. _______________________ .. code-block:: none - radial_velocity_field_name = "VT"; + radial_velocity_field_name = "VR"; radial_velocity_long_field_name = "Radial Velocity"; -The **radial_velocity_field_name** and **radial_velocity_long_field_name** parameters define the field names to give the output radial velocity grid in the netCDF output file. The parameters are used only if **compute_radial_and_radial_winds** is set to TRUE. +The **radial_velocity_field_name** and **radial_velocity_long_field_name** parameters define the field names to give the output radial velocity grid in the NetCDF output file. The parameters are used only if **compute_tangential_and_radial_winds** is set to TRUE. tc_rmw Output File @@ -199,7 +199,7 @@ The NetCDF output file contains the following dimensions: 1. *track_point* - the track points corresponding to the model output valid times -2. *pressure* - if any pressure levels are specified in the data variable list, they will be sorted and combined into a 3D NetCDF variable, which pressure as the vertical dimension and range and azimuth as the horizontal dimensions +2. *pressure* - if any pressure levels are specified in the data variable list, they will be sorted and combined into a 3D NetCDF variable, with pressure as the vertical dimension and range and azimuth as the horizontal dimensions 3. *range* - the radial dimension of the range-azimuth grid diff --git a/docs/Users_Guide/tc-stat.rst b/docs/Users_Guide/tc-stat.rst index 48dc4bde11..9e334a299c 100644 --- a/docs/Users_Guide/tc-stat.rst +++ b/docs/Users_Guide/tc-stat.rst @@ -7,7 +7,7 @@ TC-Stat Tool Introduction ============ -The TC-Stat tool ties together results from the TC-Pairs tool by providing summary statistics and filtering jobs on TCST output files. The TC-Stat tool requires TCST output from the TC-Pairs tool. See :numref:`tc_stat-output` of this user's guide for information on the TCST output format of the TC-Pairs tool. The TC-Stat tool supports several analysis job types. The **filter** job stratifies the TCST data using various conditions and thresholds described in :numref:`tc_stat-configuration-file`. The **summary** job produces summary statistics including frequency of superior performance, time-series independence calculations, and confidence intervals on the mean. The **rirw** job processes TCMPR lines, identifies adeck and bdeck rapid intensification or weakening events, populates a 2x2 contingency table, and derives contingency table statistics. The **probrirw job** processes PROBRIRW lines, populates an Nx2 probabilistic contingency table, and derives probabilistic statistics. The statistical aspects are described in :numref:`Statistical-aspects`, and practical use information for the TC-Stat tool is described in :numref:`Practical-information-1`. +The TC-Stat tool ties together results from the TC-Pairs tool by providing summary statistics and filtering jobs on TCST output files. The TC-Stat tool requires TCST output from the TC-Pairs tool. See :numref:`tc_pairs-output` of this user's guide for information on the TCST output format of the TC-Pairs tool. The TC-Stat tool supports several analysis job types. The **filter** job stratifies the TCST data using various conditions and thresholds described in :numref:`tc_stat-configuration-file`. The **summary** job produces summary statistics including frequency of superior performance, time-series independence calculations, and confidence intervals on the mean. The **rirw** job processes TCMPR lines, identifies adeck and bdeck rapid intensification or weakening events, populates a 2x2 contingency table, and derives contingency table statistics. The **probrirw** job processes PROBRIRW lines, populates an Nx2 probabilistic contingency table, and derives probabilistic statistics. The statistical aspects are described in :numref:`Statistical-aspects`, and practical use information for the TC-Stat tool is described in :numref:`Practical-information-1`. .. _Statistical-aspects: @@ -28,7 +28,7 @@ The TC-Stat tool can be used to produce summary information for a single column Confidence intervals are computed for the mean of the column of data. Confidence intervals are computed using the assumption of normality for the mean. For further information on computing confidence intervals, refer to :numref:`App_D-Confidence-Intervals` of the MET user's guide. -When operating on columns, a specific column name can be listed (e.g. TK_ERR), as well as the differences of two columns (e.g. AMAX_WIND-BMAX_WIND), and the absolute difference of the column(s) (e.g. abs(AMAX_WIND-BMAX_WIND)). Additionally, several shortcuts can be applied to choose multiple columns with a single entry. Shortcut options for the -column entry are as follows: +When operating on columns, a specific column name can be listed (e.g., TK_ERR), as well as the differences of two columns (e.g., AMAX_WIND-BMAX_WIND), and the absolute difference of the column(s) (e.g., abs(AMAX_WIND-BMAX_WIND)). Additionally, several shortcuts can be applied to choose multiple columns with a single entry. Shortcut options for the -column entry are as follows: **TRACK:** track error (TK_ERR), along-track error (ALTK_ERR), and cross-track error (CRTK_ERR) @@ -45,7 +45,7 @@ The TC-Stat tool can also be used to generate frequency of superior performance Frequency of Superior Performance ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ -The frequency of superior performance (FSP) looks at multiple model forecasts (adecks), and ranks each model relative to other model performance for the column specified in the summary job. The summary job output lists the total number of cases included in the FSP, the number of cases where the model of interest is the best (e.g.: lowest track error), the number of ties between the models, and the FSP (percent). Ties are not included in the FSP percentage; therefore the percentage may not equal 100%. +The frequency of superior performance (FSP) looks at multiple model forecasts (adecks), and ranks each model relative to other model performance for the column specified in the summary job. The summary job output lists the total number of cases included in the FSP, the number of cases where the model of interest is the best (e.g., lowest track error), the number of ties between the models, and the FSP (percent). Ties are not included in the FSP percentage; therefore the percentage may not equal 100%. Time-Series Independence ^^^^^^^^^^^^^^^^^^^^^^^^ @@ -55,9 +55,9 @@ The time-series independence evaluates effective forecast separation time using Rapid Intensification/Weakening ------------------------------- -The TC-Stat tool can be used to read TCMPR lines and compare the occurrence of rapid intensification (i.e. increase in intensity) or weakening (i.e. decrease in intensity) between the adeck and bdeck. The rapid intensification or weakening is defined by the change of maximum wind speed (i.e. **AMAX_WIND** and **BMAX_WIND** columns) over a specified amount of time. Accurately forecasting large changes in intensity is a challenging problem and this job helps quantify a model's ability to do so. +The TC-Stat tool can be used to read TCMPR lines and compare the occurrence of rapid intensification (i.e., increase in intensity) or weakening (i.e., decrease in intensity) between the adeck and bdeck. The rapid intensification or weakening is defined by the change of maximum wind speed (i.e., **AMAX_WIND** and **BMAX_WIND** columns) over a specified amount of time. Accurately forecasting large changes in intensity is a challenging problem and this job helps quantify a model's ability to do so. -Users may specify several job command options to configure the behavior of this job. Using these configurable options, the TC-Stat tool analyzes paired tracks and for each track point (i.e. each TCMPR line) determines whether rapid intensification or weakening occurred. For each point in time, it uses the forecast and BEST track event occurrence to populate a 2x2 contingency table. The job may be configured to require that forecast and BEST track events occur at exactly the same time to be considered a hit. Alternatively, the job may be configured to define a hit as long as the forecast and BEST track events occurred within a configurable time window. Using this relaxed matching criteria false alarms may be considered hits and misses may be considered correct negatives as long as the adeck and bdeck events were close enough in time. Each rirw job applies a single intensity change threshold. Therefore, assessing a model's performance with rapid intensification and weakening requires that two separate jobs be run. +Users may specify several job command options to configure the behavior of this job. Using these configurable options, the TC-Stat tool analyzes paired tracks and for each track point (i.e., each TCMPR line) determines whether rapid intensification or weakening occurred. For each point in time, it uses the forecast and BEST track event occurrence to populate a 2x2 contingency table. The job may be configured to require that forecast and BEST track events occur at exactly the same time to be considered a hit. Alternatively, the job may be configured to define a hit as long as the forecast and BEST track events occurred within a configurable time window. Using this relaxed matching criteria false alarms may be considered hits and misses may be considered correct negatives as long as the adeck and bdeck events were close enough in time. Each rirw job applies a single intensity change threshold. Therefore, assessing a model's performance with rapid intensification and weakening requires that two separate jobs be run. The RIRW job supports the **-out_stat** option to write the contingency table counts and statistics to a STAT output file. @@ -66,7 +66,7 @@ Probability of Rapid Intensification The TC-Stat tool can be used to accumulate multiple PROBRIRW lines and derive probabilistic statistics summarizing performance. The PROBRIRW line contains a probabilistic forecast for a specified intensity change along with the actual intensity change that occurred in the BEST track. Accurately forecasting the likelihood of large changes in intensity is a challenging problem and this job helps quantify a model's ability to do so. -Users may specify several job command options to configure the behavior of this job. The TC-Stat tools reads the input PROBI lines, applies the configurable options to extract a forecast probability value and BEST track event, and bins those probabilistic pairs into an Nx2 contingency table. This job writes up to four probabilistic output line types summarizing the performance. +Users may specify several job command options to configure the behavior of this job. The TC-Stat tool reads the input PROBI lines, applies the configurable options to extract a forecast probability value and BEST track event, and bins those probabilistic pairs into an Nx2 contingency table. This job writes up to four probabilistic output line types summarizing the performance. .. _Practical-information-1: @@ -96,7 +96,7 @@ The usage statement for the TC-Stat tool includes the "job" term, which refers t Required Arguments for tc_stat ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ -1. The **-lookin source** argument indicates the location of the input TCST files generated from tc_pairs. This argument can be used one or more times to specify the name of a TCST file or top-level directory containing TCST files to be processed. Multiple tcst files may be specified by using a wild card (\*). +1. The **-lookin source** argument indicates the location of the input TCST files generated from tc_pairs. This argument can be used one or more times to specify the name of a TCST file or top-level directory containing TCST files to be processed. Multiple TCST files may be specified by using a wild card (\*). 2. Either a configuration file must be specified with the **-config** option, or a **JOB COMMAND LINE** must be denoted. The **JOB COMMAND LINE** options are described in :numref:`tc_stat-configuration-file`. @@ -117,7 +117,7 @@ An example of the tc_stat calling sequence is shown below: tc_stat -lookin /home/tc_pairs/*al092010.tcst -config TCStatConfig -In this example, the TC-Stat tool uses any TCST file (output from tc_pairs) in the listed directory for the 9th Atlantic Basin storm in 2010. Filtering options and aggregated statistics are generated following configuration options specified in the **TCStatConfig** file. Further, using flags (e.g. **-basin, -column, -storm_name,** etc...) option within the job command lines may further refine these selections. See :numref:`tc_stat-configuration-file` for options available for the job command line and :numref:`config_options_tc` for how to use them. +In this example, the TC-Stat tool uses any TCST file (output from tc_pairs) in the listed directory for the 9th Atlantic Basin storm in 2010. Filtering options and aggregated statistics are generated following configuration options specified in the **TCStatConfig** file. Further, using flags (e.g., **-basin, -column, -storm_name,** etc...) option within the job command lines may further refine these selections. See :numref:`tc_stat-configuration-file` for options available for the job command line and :numref:`config_options_tc` for how to use them. .. _tc_stat-configuration-file: @@ -162,7 +162,7 @@ _________________________ amodel = []; bmodel = []; -The **amodel** and **bmodel** fields stratify by the amodel and bmodel columns based on a comma-separated list of model names used for all analysis performed. The names must be in double quotation marks (e.g.: "HWFI"). The **amodel** list specifies the model to be verified against the listed bmodel. The **bmodel** specifies the reference dataset, generally the BEST track analysis. Using the **-amodel** and **-bmodel** options within the job command lines may further refine these selections. +The **amodel** and **bmodel** fields stratify by the amodel and bmodel columns based on a comma-separated list of model names used for all analysis performed. The names must be in double quotation marks (e.g., "HWFI"). The **amodel** list specifies the model to be verified against the listed bmodel. The **bmodel** specifies the reference dataset, generally the BEST track analysis. Using the **-amodel** and **-bmodel** options within the job command lines may further refine these selections. _________________________ @@ -229,7 +229,7 @@ Option 4. A forecast is verified when a watch/warning is NOT in effect column_str_name = ["WATCH_WARN"]; column_str_val = ["NA"]; -Further information on the **column_str** and **init_str** fields are described below. Listing a comma-separated list of watch/warning types in the **column_str_val** field will be stratified by a single or multiple types of warnings. +Further information on the **column_str** and **init_str** fields is described below. Listing a comma-separated list of watch/warning types in the **column_str_val** field will be stratified by a single or multiple types of warnings. _________________________ @@ -247,7 +247,7 @@ _________________________ column_str_name = []; column_str_val = []; -The **column_str_name** and **column_str_val** fields stratify by performing string matching on non-numeric data columns. Specify a comma-separated list of columns names and values to be **included** in the analysis. The length of the **column_str_val** should match that of the **column_str_name**. Using the **-column_str name value** option within the job command lines may further refine these selections. +The **column_str_name** and **column_str_val** fields stratify by performing string matching on non-numeric data columns. Specify a comma-separated list of column names and values to be **included** in the analysis. The length of the **column_str_val** should match that of the **column_str_name**. Using the **-column_str name value** option within the job command lines may further refine these selections. _________________________ @@ -256,7 +256,7 @@ _________________________ column_str_exc_name = []; column_str_exc_val = []; -The **column_str_exc_name** and **column_str_exc_val** fields stratify by performing string matching on non-numeric data columns. Specify a comma-separated list of columns names and values to be **excluded** from the analysis. The length of the **column_str_exc_val** should match that of the **column_str_exc_name**. Using the **-column_str_exc name value** option within the job command lines may further refine these selections. +The **column_str_exc_name** and **column_str_exc_val** fields stratify by performing string matching on non-numeric data columns. Specify a comma-separated list of column names and values to be **excluded** from the analysis. The length of the **column_str_exc_val** should match that of the **column_str_exc_name**. Using the **-column_str_exc name value** option within the job command lines may further refine these selections. _________________________ @@ -334,7 +334,7 @@ _________________________ landfall_beg = "-24"; landfall_end = "00"; -The **landfall, landfall_beg**, and **landfall_end** fields specify whether only those track points occurring near landfall should be retained. The landfall retention window is defined as the hours offset from the time of landfall. Landfall is defined as the last bmodel track point before the distance to land switches from water to land. When **landfall_end** is set to zero, the track is retained from the **landfall_beg** to the time of landfall. Using the **-landfall_window** option with the job command lines may further refine these selections. The **-landfall_window** job command option takes one or two arguments in HH[MMSS] format. Use one argument to define a symmetric time window. For example, **-landfall_window 06** defines the time window +/- six hours around the landfall time. Use two arguments to define an asymmetric time window. For example, **-landfall_window 00 12** defines the time window from the landfall event to twelve hours after. +The **landfall, landfall_beg**, and **landfall_end** fields specify whether only those track points occurring near landfall should be retained. The landfall retention window is defined as the hours offset from the time of landfall. Landfall is defined as the last bmodel track point before the distance to land switches from water to land. When **landfall_end** is set to zero, the track is retained from the **landfall_beg** to the time of landfall. Using the **-landfall_window** option within the job command lines may further refine these selections. The **-landfall_window** job command option takes one or two arguments in HH[MMSS] format. Use one argument to define a symmetric time window. For example, **-landfall_window 06** defines the time window +/- six hours around the landfall time. Use two arguments to define an asymmetric time window. For example, **-landfall_window 00 12** defines the time window from the landfall event to twelve hours after. _________________________ @@ -400,7 +400,7 @@ The output generated from the TC-Stat tool contains statistics produced by the a This job command finds and filters TCST lines down to those meeting the criteria selected by the filter's options. The filtered TCST lines are written to a file specified by the **-dump_row** option. The TCST output from this job follows the TCST output description in :numref:`tc-dland` and :numref:`tc-pairs`. - The "-set_hdr" job command option can be used to override any of the output header strings (e.g. "-set_hdr DESC EVENT_EQUAL" sets the output DESC column to "EVENT_EQUAL"). +The "-set_hdr" job command option can be used to override any of the output header strings (e.g., "-set_hdr DESC EVENT_EQUAL" sets the output DESC column to "EVENT_EQUAL"). **Job: Summary** @@ -415,7 +415,7 @@ This job produces summary statistics for the column name specified by the **-col 3. "SUMMARY", which is followed by the total, mean (with confidence intervals), standard deviation, minimum value, percentiles (10th, 25th, 50th, 75th, 90th), maximum value, interquartile range, range, sum, time to independence, and frequency of superior performance. -The output columns are shown below in :numref:`table_columnar_output_summary_tc_stat` The **-by** option can also be used one or more times to make this job more powerful. Rather than running the specified job once, it will be run once for each unique combination of the entries found in the column(s) specified with the **-by** option. +The output columns are shown below in :numref:`table_columnar_output_summary_tc_stat`. The **-by** option can also be used one or more times to make this job more powerful. Rather than running the specified job once, it will be run once for each unique combination of the entries found in the column(s) specified with the **-by** option. .. _table_columnar_output_summary_tc_stat: @@ -490,16 +490,16 @@ When using multiple "-by" options, use "CASE" to reference the full case informa The PROBRIRW job produces probabilistic contingency table counts and statistics defined by placing forecast probabilities and BEST track rapid intensification events into an Nx2 contingency table. Users may specify several job command options to configure the behavior of this job: -• The **-prob_thresh n** option is required and defines which probability threshold should be evaluated. It determines which **PROB_i** column from the PROBRIRW line type is selected for the job. For example, use **-prob_thresh 30** to evaluate forecast probabilities of a 30 kt increase or use **-prob_thresh -30** to evaluate forecast probabilities of a 30 kt decrease in intensity. The default is a 30 kt increase. +• The **-probrirw_thresh n** option is required and defines which probability threshold should be evaluated. It determines which **PROB_i** column from the PROBRIRW line type is selected for the job. For example, use **-probrirw_thresh 30** to evaluate forecast probabilities of a 30 kt increase or use **-probrirw_thresh -30** to evaluate forecast probabilities of a 30 kt decrease in intensity. There is no default value. -• The **-prob_exact bool** option is a boolean defining whether the exact or maximum BEST track intensity change over the time window should be used. If true, the values in the **BDELTA** column are used. If false, the values in the **BDELTA_MAX** column are used. The default is true. +• The **-probrirw_exact bool** option is a boolean defining whether the exact or maximum BEST track intensity change over the time window should be used. If true, the values in the **BDELTA** column are used. If false, the values in the **BDELTA_MAX** column are used. The default is false. -• The **-probrirw_bdelta_thresh** threshold option defines the BEST track intensity change event threshold. This should typically be set consistent with the probability threshold (**-prob_thresh**) chosen above. The default is greater than or equal to 30 kts. +• The **-probrirw_bdelta_thresh** threshold option defines the BEST track intensity change event threshold. This should typically be set consistent with the probability threshold (**-probrirw_thresh**) chosen above. The default is greater than or equal to 30 kts. • The **-probrirw_prob_thresh threshold_list** option defines the probability thresholds used to create the output Nx2 contingency table. The default is probability bins of width 0.1. These probabilities may be specified as a list (>0.00,>0.25,>0.50,>0.75,>1.00) or using shorthand notation (==0.25) for bins of equal width. • The **-out_line_type** option defines the output data that should be written. This job can write PCT, PSTD, PJC, and PRC output line types. The default is PCT and PSTD. Please see :numref:`table_PS_format_info_PCT` through :numref:`table_PS_format_info_PRC` for more details. -Users may also specify the **-out_alpha** option to define the alpha value for the confidence intervals in the PSTD output line type. Multiple values in the **RI_WINDOW** column cannot be combined in a single PROBRIRW job since the BEST track intensity threshold should change for each. Using the **-by RI_WINDOW** option or **-column_thresh RI_WINDOW ==24** option provide convenient ways avoiding this problem. +Users may also specify the **-out_alpha** option to define the alpha value for the confidence intervals in the PSTD output line type. Multiple values in the **RI_WINDOW** column cannot be combined in a single PROBRIRW job since the BEST track intensity threshold should change for each. Using the **-by RI_WINDOW** option or **-column_thresh RI_WINDOW ==24** option provide convenient ways of avoiding this problem. -Users should note that for the PROBRIRW line type, **PROBRI_PROB** is a derived column name. The -probrirw_thresh option defines the probabilities of interest (e.g. **-probrirw_thresh 30**) and the **PROBRI_PROB** column name refers to those probability values, regardless of their column number. For example, the job command options **-probrirw_thresh 30 -column_thresh PROBRI_PROB >0** select 30 kt probabilities and match probability values greater than 0. +Users should note that for the PROBRIRW line type, **PROBRI_PROB** is a derived column name. The -probrirw_thresh option defines the probabilities of interest (e.g., **-probrirw_thresh 30**) and the **PROBRI_PROB** column name refers to those probability values, regardless of their column number. For example, the job command options **-probrirw_thresh 30 -column_thresh PROBRI_PROB >0** select 30 kt probabilities and match probability values greater than 0. diff --git a/docs/Users_Guide/wavelet-stat.rst b/docs/Users_Guide/wavelet-stat.rst index f2fc74594d..961c357430 100644 --- a/docs/Users_Guide/wavelet-stat.rst +++ b/docs/Users_Guide/wavelet-stat.rst @@ -15,9 +15,9 @@ The Intensity-Scale technique is one of the recently developed verification appr The Intensity-Scale verification technique, as most of the spatial verification approaches, compares a forecast field to an observation field. To apply the Intensity-Scale verification approach, observations need to be defined over the same spatial domain of the forecast to be verified. -Within the spatial verification approaches, the Intensity-Scale technique belongs to the scale-decomposition (or scale-separation) verification approaches. The scale-decomposition approaches enable users to perform the verification on different spatial scales. Weather phenomena on different scales (e.g. frontal systems versus convective showers) are often driven by different physical processes. Verification on different spatial scales can therefore provide deeper insights into model performance at simulating these different processes. +Within the spatial verification approaches, the Intensity-Scale technique belongs to the scale-decomposition (or scale-separation) verification approaches. The scale-decomposition approaches enable users to perform the verification on different spatial scales. Weather phenomena on different scales (e.g., frontal systems versus convective showers) are often driven by different physical processes. Verification on different spatial scales can therefore provide deeper insights into model performance at simulating these different processes. -The spatial scale components are obtained usually by applying a single band spatial filter to the forecast and observation fields (e.g. Fourier, Wavelets). The scale-decomposition approaches measure error, bias and skill of the forecast on each different scale component. The scale-decomposition approaches therefore provide feedback on the scale dependency of the error and skill, on the no-skill to skill transition scale, and on the capability of the forecast of reproducing the observed scale structure. +The spatial scale components are obtained usually by applying a single band spatial filter to the forecast and observation fields (e.g., Fourier, Wavelets). The scale-decomposition approaches measure error, bias and skill of the forecast on each different scale component. The scale-decomposition approaches therefore provide feedback on the scale dependency of the error and skill, on the no-skill to skill transition scale, and on the capability of the forecast of reproducing the observed scale structure. The Intensity-Scale technique evaluates the forecast skill as a function of the intensity values and of the spatial scale of the error. The scale components are obtained by applying a two dimensional Haar wavelet filter. Note that wavelets, because of their locality, are suitable for representing discontinuous fields characterized by few sparse non-zero features, such as precipitation. Moreover, the technique is based on a categorical approach, which is a robust and resistant approach, suitable for non-normally distributed variables, such as precipitation. The intensity-scale technique was specifically designed to cope with the difficult characteristics of precipitation fields, and for the verification of spatial precipitation forecasts. However, the intensity-scale technique can also be applied to verify other variables, such as cloud fraction. @@ -35,19 +35,19 @@ The Intensity Scale approach can be summarized in the following 5 steps: 2. The binary forecast and observation fields obtained from the thresholding are then decomposed into the sum of components on different scales, by using a 2D Haar wavelet filter (:numref:`wavelet-stat_NIMROD_diff`). Note that the scale components are fields, and their sum adds up to the original binary field. For a forecast defined over square domain of :math:`{2^n} \times {2^n}` grid-points, the scale components are **n+1: n** mother wavelet components + the largest father wavelet (or scale-function) component. The **n** mother wavelet components have resolution equal to **1, 2, 4, ...** :math:`{2^{n-1}}` grid-points. The largest father wavelet component is a constant field over the :math:`{2^n} \times {2^n}` grid-point domain with value equal to the field mean. -**Note** that the wavelet transform is a linear operator: this implies that the difference of the spatial scale components of the binary forecast and observation fields (:numref:`wavelet-stat_NIMROD_diff`) are equal to the spatial scale components of the difference of the binary forecast and observation fields (:numref:`wavelet-stat_NIMROD_binary_fcst_and_obs`), and these scale components also add up to the original binary field difference (:numref:`wavelet-stat_NIMROD_3h_fcst`). The intensity-scale technique considers thus the spatial scale of the error. For the case illustrated (:numref:`wavelet-stat_NIMROD_3h_fcst` and :numref:`wavelet-stat_NIMROD_binary_fcst_and_obs`) note the large error associated at the scale of 160 km, due the storm, 160km displaced almost its entire length. +**Note** that the wavelet transform is a linear operator: this implies that the difference of the spatial scale components of the binary forecast and observation fields (:numref:`wavelet-stat_NIMROD_diff`) are equal to the spatial scale components of the difference of the binary forecast and observation fields (:numref:`wavelet-stat_NIMROD_binary_fcst_and_obs`), and these scale components also add up to the original binary field difference (:numref:`wavelet-stat_NIMROD_3h_fcst`). The intensity-scale technique considers thus the spatial scale of the error. For the case illustrated (:numref:`wavelet-stat_NIMROD_3h_fcst` and :numref:`wavelet-stat_NIMROD_binary_fcst_and_obs`) note the large error associated at the scale of 160 km, due to the storm, 160km displaced almost its entire length. -**Note** also that the means of the binary forecast and observation fields (i.e. their largest father wavelet components) are equal to the proportion of forecast and observed events above the threshold, **(a+b)/n** and **(a+c)/n**, evaluated from the contingency table counts (:numref:`contingency_table_counts`) obtained from the original forecast and observation fields by thresholding with the same threshold used to obtain the binary forecast and observation fields. This relation is intuitive when observing forecast and observation binary fields and their corresponding contingency table image (:numref:`wavelet-stat_NIMROD_3h_fcst`). The comparison of the largest father wavelet component of binary forecast and observation fields therefore provides feedback on the whole field bias. +**Note** also that the means of the binary forecast and observation fields (i.e., their largest father wavelet components) are equal to the proportion of forecast and observed events above the threshold, **(a+b)/n** and **(a+c)/n**, evaluated from the contingency table counts (:numref:`contingency_table_counts`) obtained from the original forecast and observation fields by thresholding with the same threshold used to obtain the binary forecast and observation fields. This relation is intuitive when observing forecast and observation binary fields and their corresponding contingency table image (:numref:`wavelet-stat_NIMROD_3h_fcst`). The comparison of the largest father wavelet component of binary forecast and observation fields therefore provides feedback on the whole field bias. -3. For each threshold (**t**) and for each scale component (**j**) of the binary forecast and observation, the Mean Squared Error (MSE) is then evaluated (:numref:`wavelet-stat_MSE_percent_NIMROD`). The error is usually large for small thresholds, and decreases as the threshold increases. This behavior is partially artificial, and occurs because the smaller the threshold the more events will exceed it, and therefore the larger would be the error, since the error tends to be proportional to the amount of events in the binary fields. The artificial effect can be diminished by normalization: because of the wavelet orthogonal properties, the sum of the MSE of the scale components is equal to the MSE of the original binary fields: :math:`MSE(t) = j MSE(t,j)`. Therefore, the percentage that the MSE for each scale contributes to the total MSE may be computed: for a given threshold, **t**, :math:`{MSE\%}(t,j) = {MSE}(t,j)/ {MSE}(t)`. The MSE% does not exhibit the threshold dependency, and usually shows small errors on large scales and large errors on small scales, with the largest error associated to the smallest scale and highest threshold. For the NIMROD case illustrated, note the large error at 160 km and between the thresholds of and 4 mm/h, due to the storm, 160km displaced almost its entire length. +3. For each threshold (**t**) and for each scale component (**j**) of the binary forecast and observation, the Mean Squared Error (MSE) is then evaluated (:numref:`wavelet-stat_MSE_percent_NIMROD`). The error is usually large for small thresholds, and decreases as the threshold increases. This behavior is partially artificial, and occurs because the smaller the threshold the more events will exceed it, and therefore the larger would be the error, since the error tends to be proportional to the amount of events in the binary fields. The artificial effect can be diminished by normalization: because of the wavelet orthogonal properties, the sum of the MSE of the scale components is equal to the MSE of the original binary fields: :math:`MSE(t) = \sum_j MSE(t,j)`. Therefore, the percentage that the MSE for each scale contributes to the total MSE may be computed: for a given threshold, **t**, :math:`{MSE\%}(t,j) = {MSE}(t,j)/ {MSE}(t)`. The MSE% does not exhibit the threshold dependency, and usually shows small errors on large scales and large errors on small scales, with the largest error associated to the smallest scale and highest threshold. For the NIMROD case illustrated, note the large error at 160 km and between the thresholds of ½ and 4 mm/h, due to the storm, 160km displaced almost its entire length. **Note** that the MSE of the original binary fields is equal to the proportion of the counts of misses (**c/n**) and false alarms (**b/n**) for the contingency table (:numref:`contingency_table_counts`) obtained from the original forecast and observation fields by thresholding with the same threshold used to obtain the binary forecast and observation fields: :math:`{MSE}(t)=(b+c)/n`. This relation is intuitive when comparing the forecast and observation binary field difference and their corresponding contingency table image (:numref:`contingency_table_counts`). -4. The MSE for the random binary forecast and observation fields is estimated by :math:`{MSE}(t) {random}= {FBI}*{Br}*(1-{Br}) + {Br}*(1- {FBI}*{Br})`, where :math:`{FBI}=(a+b)/(a+c)` is the frequency bias index and :math:`{Br}=(a+c)/n` is the sample climatology from the contingency table (:numref:`contingency_table_counts`) obtained from the original forecast and observation fields by thresholding with the same threshold used to obtain the binary forecast and observation fields. This formula follows by considering the :ref:`Murphy and Winkler (1987) ` framework, applying the Bayes' theorem to express the joint probabilities **b/n** and **c/n** as product of the marginal and conditional probability (e.g. :ref:`Jolliffe and Stephenson, 2012 `; :ref:`Wilks, 2010 `), and then noticing that for a random forecast the conditional probability is equal to the unconditional one, so that **b/n** and **c/n** are equal to the product of the corresponding marginal probabilities solely. +4. The MSE for the random binary forecast and observation fields is estimated by :math:`{MSE}(t) {random}= {FBI}*{Br}*(1-{Br}) + {Br}*(1- {FBI}*{Br})`, where :math:`{FBI}=(a+b)/(a+c)` is the frequency bias index and :math:`{Br}=(a+c)/n` is the sample climatology from the contingency table (:numref:`contingency_table_counts`) obtained from the original forecast and observation fields by thresholding with the same threshold used to obtain the binary forecast and observation fields. This formula follows by considering the :ref:`Murphy and Winkler (1987) ` framework, applying the Bayes' theorem to express the joint probabilities **b/n** and **c/n** as product of the marginal and conditional probability (e.g., :ref:`Jolliffe and Stephenson, 2012 `; :ref:`Wilks, 2010 `), and then noticing that for a random forecast the conditional probability is equal to the unconditional one, so that **b/n** and **c/n** are equal to the product of the corresponding marginal probabilities solely. 5. For each threshold (**t**) and scale component (**j**), the skill score based on the MSE of binary forecast and observation scale components is evaluated (:numref:`wavelet-stat_Intensity_Scale_skill_score_NIMROD`). The standard skill score definition as in :ref:`Jolliffe and Stephenson (2012) ` or :ref:`Wilks (2010) ` is used, and random chance is used as reference forecast. The MSE for the random binary forecast is equipartitioned on the **n+1** scales to evaluate the skill score: :math:`{SS} (t,j)=1- {MSE}(t,j)*(n+1)/ {MSE}(t) {random}` -The Intensity-Scale (IS) skill score evaluates the forecast skill as a function of the precipitation intensity and of the spatial scale of the error. Positive values of the IS skill score are associated with a skillful forecast, whereas negative values are associated with no skill. Usually large scales exhibit positive skill (large scale events, such as fronts, are well predicted), whereas small scales exhibit negative skill (small scale events, such as convective showers, are less predictable), and the smallest scale and highest thresholds exhibit the worst skill. For the NIMROD case illustrated note the negative skill associated with the 160 km scale, for the thresholds to 4 mm/h, due to the 160 km storm displaced almost its entire length. +The Intensity-Scale (IS) skill score evaluates the forecast skill as a function of the precipitation intensity and of the spatial scale of the error. Positive values of the IS skill score are associated with a skillful forecast, whereas negative values are associated with no skill. Usually large scales exhibit positive skill (large scale events, such as fronts, are well predicted), whereas small scales exhibit negative skill (small scale events, such as convective showers, are less predictable), and the smallest scale and highest thresholds exhibit the worst skill. For the NIMROD case illustrated note the negative skill associated with the 160 km scale, for the thresholds ½ to 4 mm/h, due to the 160 km storm displaced almost its entire length. .. _contingency_table_counts: @@ -80,31 +80,31 @@ The Intensity-Scale (IS) skill score evaluates the forecast skill as a function .. figure:: figure/wavelet-stat_NIMROD_3h_fcst.png - NIMROD 3h lead-time forecast and corresponding verifying analysis field (precipitation rate in mm/h, valid the 05/29/99 at 15:00 UTC); forecast and analysis binary fields obtained for a threshold of 1mm/h, the binary field difference has their corresponding Contingency Table Image (see :numref:`contingency_table_counts`). The forecast shows a storm of 160 km displaced almost its entire length. + NIMROD 3h lead-time forecast and corresponding verifying analysis field (precipitation rate in mm/h, valid the 05/29/99 at 15:00 UTC); forecast and analysis binary fields obtained for a threshold of 1mm/h, the binary field difference has their corresponding Contingency Table Image (see :numref:`contingency_table_counts`). The forecast shows a storm of 160 km displaced almost its entire length. .. _wavelet-stat_NIMROD_binary_fcst_and_obs: .. figure:: figure/wavelet-stat_NIMROD_binary_fcst_and_obs.png - NIMROD binary forecast (top) and binary analysis (bottom) spatial scale components obtained by a 2D Haar wavelet transform (th=1 mm/h). Scales 1 to 8 refer to mother wavelet components (5, 10, 20, 40, 80, 160, 320, 640 km resolution); scale 9 refers to the largest father wavelet component (1280 km resolution). + NIMROD binary forecast (top) and binary analysis (bottom) spatial scale components obtained by a 2D Haar wavelet transform (th=1 mm/h). Scales 1 to 8 refer to mother wavelet components (5, 10, 20, 40, 80, 160, 320, 640 km resolution); scale 9 refers to the largest father wavelet component (1280 km resolution). .. _wavelet-stat_NIMROD_diff: .. figure:: figure/wavelet-stat_NIMROD_diff.png - NIMROD binary field difference spatial scale components obtained by a 2D Haar wavelet transform (th=1 mm/h). Scales 1 to 8 refer to mother wavelet components (5, 10, 20, 40, 80, 160, 320, 640 km resolution); scale 9 refers to the largest father wavelet component (1280 km resolution). Note the large error at the scale 6 = 160 km, due to the storm, 160 km displaced almost of its entire length. + NIMROD binary field difference spatial scale components obtained by a 2D Haar wavelet transform (th=1 mm/h). Scales 1 to 8 refer to mother wavelet components (5, 10, 20, 40, 80, 160, 320, 640 km resolution); scale 9 refers to the largest father wavelet component (1280 km resolution). Note the large error at the scale 6 = 160 km, due to the storm, 160 km displaced almost its entire length. .. _wavelet-stat_MSE_percent_NIMROD: .. figure:: figure/wavelet-stat_MSE_percent_NIMROD.png - MSE and MSE % for the NIMROD binary forecast and analysis spatial scale components. In the MSE%, note the large error associated with the scale 6 = 160 km, for the thresholds ½ to 4 mm/h, associated with the displaced storm. + MSE and MSE % for the NIMROD binary forecast and analysis spatial scale components. In the MSE%, note the large error associated with the scale 6 = 160 km, for the thresholds ½ to 4 mm/h, associated with the displaced storm. .. _wavelet-stat_Intensity_Scale_skill_score_NIMROD: .. figure:: figure/wavelet-stat_Intensity_Scale_skill_score_NIMROD.png - Intensity-Scale skill score for the NIMROD forecast and analysis shown in :numref:`wavelet-stat_NIMROD_3h_fcst`. The skill score is a function of the intensity of the precipitation rate and spatial scale of the error. Note the negative skill associated with the scale 6 = 160 km, for the thresholds to 4 mm/h, associated with the displaced storm. + Intensity-Scale skill score for the NIMROD forecast and analysis shown in :numref:`wavelet-stat_NIMROD_3h_fcst`. The skill score is a function of the intensity of the precipitation rate and spatial scale of the error. Note the negative skill associated with the scale 6 = 160 km, for the thresholds ½ to 4 mm/h, associated with the displaced storm. @@ -114,11 +114,11 @@ In addition to the MSE and the SS, the energy squared is also evaluated, for eac .. figure:: figure/wavelet-stat_energy_squared_NIMROD.png - Energy squared and energy squared percentages, for each threshold and sale, for the NIMROD forecast and analysis, and forecast and analysis En2 and En2% relative differences. + Energy squared and energy squared percentages, for each threshold and scale, for the NIMROD forecast and analysis, and forecast and analysis En2 and En2% relative differences. -The En2 bias for each threshold and scale is assessed by the En2 relative difference, equal to the difference between forecast and observed squared energies normalized by their sum: :math:`{En2}(F)- {En2}(O)]/[{En2}(F)+ {En2}(O)]`. Since defined in such a fashion, the En2 relative difference accounts for the difference between forecast and observation squared energies relative to their magnitude, and it is sensitive therefore to the ratio of the forecast and observed squared energies. The En2 relative difference ranges between -1 and 1, positive values indicate over-forecast and negative values indicate under-forecast. For the NIMROD case illustrated the forecast exhibits over-forecast for small thresholds, quite pronounced on the large scales, and under-forecast for high thresholds. +The En2 bias for each threshold and scale is assessed by the En2 relative difference, equal to the difference between forecast and observed squared energies normalized by their sum: :math:`[{En2}(F)- {En2}(O)]/[{En2}(F)+ {En2}(O)]`. Since defined in such a fashion, the En2 relative difference accounts for the difference between forecast and observation squared energies relative to their magnitude, and it is sensitive therefore to the ratio of the forecast and observed squared energies. The En2 relative difference ranges between -1 and 1, positive values indicate over-forecast and negative values indicate under-forecast. For the NIMROD case illustrated the forecast exhibits over-forecast for small thresholds, quite pronounced on the large scales, and under-forecast for high thresholds. -As for the MSE, the sum of the energy of the scale components is equal to the energy of the original binary field: :math:`{En2}(t) = j \ {En2}(t,j)`. Therefore, the percentage that the En2 for each scale contributes the total En2 may be computed: for a given threshold, **t**, :math:`{En2\%}(t,j) = {En2}(t,j)/ {En2}(t)`. Usually, for precipitation fields, low thresholds exhibit most of the energy percentage on large scales (and less percentage on the small scales), since low thresholds are associated with large scale features, such as fronts. On the other hand, for higher thresholds, the energy percentage is usually larger on small scales, since intense events are associated with small scales features, such as convective cells or showers. The comparison of the forecast and observation squared energy percentages provides feedback on how the events are distributed across the scales, and enables the comparison of forecast and observation scale structure. +As for the MSE, the sum of the energy of the scale components is equal to the energy of the original binary field: :math:`{En2}(t) = \sum_j {En2}(t,j)`. Therefore, the percentage that the En2 for each scale contributes the total En2 may be computed: for a given threshold, **t**, :math:`{En2\%}(t,j) = {En2}(t,j)/ {En2}(t)`. Usually, for precipitation fields, low thresholds exhibit most of the energy percentage on large scales (and less percentage on the small scales), since low thresholds are associated with large scale features, such as fronts. On the other hand, for higher thresholds, the energy percentage is usually larger on small scales, since intense events are associated with small scales features, such as convective cells or showers. The comparison of the forecast and observation squared energy percentages provides feedback on how the events are distributed across the scales, and enables the comparison of forecast and observation scale structure. For the NIMROD case illustrated, the scale structure is assessed again by the relative difference, but calculated of the squared energy percentages. For small thresholds the forecast overestimates the number of large scale events and underestimates the number of small scale events, in proportion to the total number of events. On the other hand, for larger thresholds the forecast underestimates the number of large scale events and overestimates the number of small scale events, again in proportion to the total number of events. Overall it appears that the forecast overestimates the percentage of events associated with high occurrence, and underestimates the percentage of events associated with low occurrence. The En2% for the 64 mm/h thresholds is homogeneously underestimated for all the scales, since the forecast does not have any event exceeding this threshold. @@ -129,25 +129,25 @@ Note that the energy squared of the observation binary field is identical to the The Spatial Domain Constraints ------------------------------ -The Intensity-Scale technique is constrained by the fact that orthogonal wavelets (discrete wavelet transforms) are usually performed dyadic domains, square domains of :math:`{2^n} \times {2^n}` grid-points. The Wavelet-Stat tool handles this issue based on settings in the configuration file by defining tiles of dimensions :math:`{2^n} \times {2^n}` over the input domain in the following ways: +The Intensity-Scale technique is constrained by the fact that orthogonal wavelets (discrete wavelet transforms) are usually performed on dyadic domains, square domains of :math:`{2^n} \times {2^n}` grid-points. The Wavelet-Stat tool handles this issue based on settings in the configuration file by defining tiles of dimensions :math:`{2^n} \times {2^n}` over the input domain in the following ways: 1. User-Defined Tiling: The user may define one or more tiles of size :math:`{2^n} \times {2^n}` over their domain to be applied. This is done by selecting the grid coordinates for the lower-left corner of the tile(s) and the tile dimension to be used. If the user specifies more than one tile, the Intensity-Scale method will be applied to each tile separately. At the end, the results will automatically be aggregated across all the tiles and written out with the results for each of the individual tiles. Users are encouraged to select tiles which consist entirely of valid data. 2. Automated Tiling: This tiling method is essentially the same as the user-defined tiling method listed above except that the tool automatically selects the location and size of the tile(s) to be applied. It figures out the maximum tile of dimension :math:`{2^n} \times {2^n}` that fits within the domain and places the tile at the center of the domain. For domains that are very elongated in one direction, it defines as many of these tiles as possible that fit within the domain. -3. Padding: If the domain size is only slightly smaller than :math:`{2^n} \times {2^n}`, for certain variables (e.g. precipitation), it is advisable to expand the domain out to :math:`{2^n} \times {2^n}` grid-points by adding extra rows and/or columns of fill data. For precipitation variables, a fill value of zero is used. For continuous variables, such as temperature, the fill value is defined as the mean of the valid data in the rest of the field. A drawback to the padding method is the introduction of artificial data into the original field. Padding should only be used when a very small number of rows and/or columns need to be added. +3. Padding: If the domain size is only slightly smaller than :math:`{2^n} \times {2^n}`, for certain variables (e.g., precipitation), it is advisable to expand the domain out to :math:`{2^n} \times {2^n}` grid-points by adding extra rows and/or columns of fill data. For precipitation variables, a fill value of zero is used. For continuous variables, such as temperature, the fill value is defined as the mean of the valid data in the rest of the field. A drawback to the padding method is the introduction of artificial data into the original field. Padding should only be used when a very small number of rows and/or columns need to be added. Aggregation of Statistics on Multiple Cases ------------------------------------------- -The Stat-Analysis tool aggregates the intensity scale technique results. Since the results are scale-dependent, it is sensible to aggregate results from multiple model runs (e.g. daily runs for a season) on the same spatial domain, so that the scale components for each singular case will be the same number, and the domain, if not a square domain of :math:`{2^n} \times {2^n}` grid-points, will be treated in the same fashion. Similarly, the intensity thresholds for each run should all be the same. +The Stat-Analysis tool aggregates the intensity scale technique results. Since the results are scale-dependent, it is sensible to aggregate results from multiple model runs (e.g., daily runs for a season) on the same spatial domain, so that the scale components for each singular case will be the same number, and the domain, if not a square domain of :math:`{2^n} \times {2^n}` grid-points, will be treated in the same fashion. Similarly, the intensity thresholds for each run should all be the same. The MSE and forecast and observation squared energy for each scale and thresholds are aggregated simply with a weighted average, where weights are proportional to the number of grid-points used in each single run to evaluate the statistics. If the same domain is always used (and it should) the weights result all the same, and the weighted averaging is a simple mean. For each threshold, the aggregated Br is equal to the aggregated squared energy of the binary observation field, and the aggregated FBI is obtained as the ratio of the aggregated squared energies of the forecast and observation binary fields. From aggregated Br and FBI, the MSErandom for the aggregated runs can be evaluated using the same formula as for the single run. Finally, the Intensity-Scale Skill Score is evaluated by using the aggregated statistics within the same formula used for the single case. Practical Information ===================== -The following sections describe the usage statement, required arguments and optional arguments for the Stat-Analysis tool. +The following sections describe the usage statement, required arguments and optional arguments for the Wavelet-Stat tool. wavelet_stat Usage ------------------ @@ -205,7 +205,7 @@ wavelet_stat Configuration File The default configuration file for the Wavelet-Stat tool, **WaveletStatConfig_default**, can be found in the installed *share/met/config* directory. Another version of the configuration file is provided in *scripts/config*. We recommend that users make a copy of the default (or other) configuration file prior to modifying it. The contents are described in more detail below. -Note that environment variables may be used when editing configuration files, as described in the :numref:`config_env_vars`. +Note that environment variables may be used when editing configuration files, as described in :numref:`config_env_vars`. _______________________ @@ -263,7 +263,7 @@ The **grid_decomp_flag** variable specifies how tiling should be performed: • **TILE** indicates that the user-defined tiles should be applied. -• **PAD** indicated that the data should be padded out to the nearest dimension of :math:`{2^n} \times {2^n}` +• **PAD** indicates that the data should be padded out to the nearest dimension of :math:`{2^n} \times {2^n}` The **width** and **location** variables allow users to manually define the tiles of dimension they would like to apply. The x_ll and y_ll variables specify the location of one or more lower-left tile grid (x, y) points. @@ -276,7 +276,7 @@ _______________________ member = 2; } -The **wavelet_flag** and **wavelet_k** variables specify the type and shape of the wavelet to be used for the scale decomposition. The :ref:`Casati et al. (2004) ` method uses a Haar wavelet which is a good choice for discontinuous fields like precipitation. However, users may choose to apply any wavelet family/shape that is available in the GNU Scientific Library. Values for the **wavelet_flag** variable, and associated choices for k, are described below: +The **wavelet.type** and **wavelet.member** entries specify the type and shape of the wavelet to be used for the scale decomposition. The :ref:`Casati et al. (2004) ` method uses a Haar wavelet which is a good choice for discontinuous fields like precipitation. However, users may choose to apply any wavelet family/shape that is available in the GNU Scientific Library. Values for the **type** entry, and associated choices for **member**, are described below: • **HAAR** for the Haar wavelet (member = 2). @@ -326,7 +326,7 @@ The output ASCII files are named similarly: wavelet_stat_PREFIX_HHMMSSL_YYYYMMDD_HHMMSSV_TYPE.txt where TYPE is isc to indicate that this is an intensity-scale line type. -The format of the STAT and ASCII output of the Wavelet-Stat tool is similar to the format of the STAT and ASCII output of the Point-Stat tool. Please refer to the tables in :numref:`point_stat-output` for a description of the common output for STAT files types. The information contained in the STAT and isc files are identical. However, for consistency with the STAT files produced by other tools, the STAT file will only have names for the header columns. The isc file contains names for all columns. The format of the ISC line type is explained in the following table. +The format of the STAT and ASCII output of the Wavelet-Stat tool is similar to the format of the STAT and ASCII output of the Point-Stat tool. Please refer to the tables in :numref:`point_stat-output` for a description of the common output for STAT file types. The information contained in the STAT and isc files are identical. However, for consistency with the STAT files produced by other tools, the STAT file will only have names for the header columns. The isc file contains names for all columns. The format of the ISC line type is explained in the following table. .. _table_WS_header_info_ws_outputs: diff --git a/docs/_static/theme_override.css b/docs/_static/theme_override.css index a03592c4b9..2a3a87c0fa 100644 --- a/docs/_static/theme_override.css +++ b/docs/_static/theme_override.css @@ -10,19 +10,19 @@ div[class^="highlight"] a:hover { /* Background highlight color (white) */ div[class^="highlight"] { - background-color: #FFFFFF; + background-color: #FFFFFF; } /* override table width restrictions */ @media screen and (min-width: 767px) { .wy-table-responsive table td { - /* !important prevents the common CSS stylesheets from overriding - this as on RTD they are loaded after this stylesheet */ - white-space: normal !important; + /* !important prevents the common CSS stylesheets from overriding + this as on RTD they are loaded after this stylesheet */ + white-space: normal !important; } .wy-table-responsive { - overflow: visible !important; + overflow: visible !important; } } diff --git a/docs/conf.py b/docs/conf.py index 0a36069e19..2a59f0b3a8 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -42,24 +42,24 @@ latex_master_doc = 'Users_Guide/index' latex_elements = { - # The paper size ('letterpaper' or 'a4paper'). - # - 'papersize': 'letterpaper', - 'releasename':"{version}", - 'fncychap': '\\usepackage{fncychap}', - 'fontpkg': '\\usepackage{amsmath,amsfonts,amssymb,amsthm,float}', - 'inputenc': '\\usepackage[utf8]{inputenc}', - 'fontenc': '\\usepackage[LGR,T1]{fontenc}', + # The paper size ('letterpaper' or 'a4paper'). + # + 'papersize': 'letterpaper', + 'releasename':"{version}", + 'fncychap': '\\usepackage{fncychap}', + 'fontpkg': '\\usepackage{amsmath,amsfonts,amssymb,amsthm,float}', + 'inputenc': '\\usepackage[utf8]{inputenc}', + 'fontenc': '\\usepackage[LGR,T1]{fontenc}', - 'figure_align':'H', - 'pointsize': '11pt', + 'figure_align':'H', + 'pointsize': '11pt', - 'preamble': r''' - \usepackage{charter} - \usepackage[defaultsans]{lato} - \usepackage{inconsolata} - \setcounter{secnumdepth}{4} - \setcounter{tocdepth}{4} + 'preamble': r''' + \usepackage{charter} + \usepackage[defaultsans]{lato} + \usepackage{inconsolata} + \setcounter{secnumdepth}{4} + \setcounter{tocdepth}{4} ''', 'sphinxsetup': \ @@ -69,9 +69,9 @@ HeaderFamily=\\rmfamily\\bfseries, \ InnerLinkColor={rgb}{0,0,1}, \ OuterLinkColor={rgb}{0,0,1}', - 'maketitle': '\\sphinxmaketitle', -# 'tableofcontents': ' ', - 'printindex': ' ' + 'maketitle': '\\sphinxmaketitle', +# 'tableofcontents': ' ', + 'printindex': ' ' } # Grouping the document tree into LaTeX files. List of tuples diff --git a/docs/index.rst b/docs/index.rst index f5da2d1057..6c1bc960ee 100644 --- a/docs/index.rst +++ b/docs/index.rst @@ -43,7 +43,7 @@ and analysis tools to provide the same primary functionality as the EMC VSDB system, and also included a spatial verification package called MODE. Over the years, MET and VSDB packages grew in complexity. Verification -capability at other NOAA laboratories, such as ESRL, were also under heavy +capability at other NOAA laboratories, such as ESRL, was also under heavy development. An effort to unify verification capability was first started under the HIWPP project and led by NOAA ESRL. In 2015, the NGGPS Program Office started working groups to focus on several aspects of the @@ -85,15 +85,15 @@ follows: components of METplus tools for statistical aggregation, event equalization, and other analysis needs * **METplotpy** - suite of Python-based scripts to plot MET output, - and in come cases provide additional post-processing of output prior + and in some cases provide additional post-processing of output prior to plotting -* **METdatadb** - database to store MET output and to be used by both +* **METdataio** - database to store MET output and to be used by both METviewer and METexpress The umbrella repository will be brought together by using a software package called `manage_externals `_ developed by the Community Earth System Modeling (CESM) team, hosted at NCAR -and NOAA Earth System's Research Laboratory. The manage_externals paackage +and NOAA Earth System Research Laboratory. The manage_externals package was developed because CESM is comprised of a number of different components that are developed and managed independently. Each component also may have additional "external" dependencies that need to be maintained independently. @@ -109,10 +109,10 @@ Acronyms * **VSDB** - Verification Statistics Data Base * **MODE** - Method for Object-Based Diagnostic Evaluation * **UFS** - Unified Forecast System -* **SIMA** -System for Integrated Modeling of the Atmosphere -* **ESRL** - Earth Systems Research Laboratory -* **HIWPP** - High Impact Weather Predication Project -* **NGGPS** - Next Generation Global Predicatio System +* **SIMA** - System for Integrated Modeling of the Atmosphere +* **ESRL** - Earth System Research Laboratory +* **HIWPP** - High Impact Weather Prediction Project +* **NGGPS** - Next Generation Global Prediction System * **GSD** - Global Systems Division Authors @@ -155,68 +155,68 @@ To cite this documentation in publications, please refer to the MET User's Guide .. rubric:: Organization .. [#NCAR] `National Center for Atmospheric Research, Research - Applications Laboratory `_, `Developmental Testbed Center `_ + Applications Laboratory `_, `Developmental Testbed Center `_ .. [#CIRA] `Cooperative Institute for Research in the Atmosphere at - National Oceanic and Atmospheric Administration (NOAA) Earth - System Research Laboratory `_ + National Oceanic and Atmospheric Administration (NOAA) Earth + System Research Laboratory `_ .. toctree:: - :hidden: - :caption: Training + :hidden: + :caption: Training - METplus Tutorial - Training Series - Featured Topics + METplus Tutorial + Training Series + Featured Topics .. toctree:: - :hidden: - :caption: METplus + :hidden: + :caption: METplus - User's Guide - Contributor's Guide - Verification Datasets Guide - Release Guide + User's Guide + Contributor's Guide + Verification Datasets Guide + Release Guide .. toctree:: - :hidden: - :caption: MET + :hidden: + :caption: MET - Users_Guide/index - Contributors_Guide/index + Users_Guide/index + Contributors_Guide/index .. toctree:: - :hidden: - :caption: METexpress + :hidden: + :caption: METexpress - User's Guide + User's Guide .. toctree:: - :hidden: - :caption: METviewer + :hidden: + :caption: METviewer - User's Guide - Contributor's Guide + User's Guide + Contributor's Guide .. toctree:: - :hidden: - :caption: METplotpy + :hidden: + :caption: METplotpy - User's Guide - Contributor's Guide + User's Guide + Contributor's Guide .. toctree:: - :hidden: - :caption: METcalcpy + :hidden: + :caption: METcalcpy - User's Guide - Contributor's Guide + User's Guide + Contributor's Guide .. toctree:: - :hidden: - :caption: METdataio + :hidden: + :caption: METdataio - User's Guide - Contributor's Guide + User's Guide + Contributor's Guide Index diff --git a/docs/make.bat b/docs/make.bat index 2119f51099..16b0638349 100644 --- a/docs/make.bat +++ b/docs/make.bat @@ -5,7 +5,7 @@ pushd %~dp0 REM Command file for Sphinx documentation if "%SPHINXBUILD%" == "" ( - set SPHINXBUILD=sphinx-build + set SPHINXBUILD=sphinx-build ) set SOURCEDIR=. set BUILDDIR=_build @@ -14,15 +14,15 @@ if "%1" == "" goto help %SPHINXBUILD% >NUL 2>NUL if errorlevel 9009 ( - echo. - echo.The 'sphinx-build' command was not found. Make sure you have Sphinx - echo.installed, then set the SPHINXBUILD environment variable to point - echo.to the full path of the 'sphinx-build' executable. Alternatively you - echo.may add the Sphinx directory to PATH. - echo. - echo.If you don't have Sphinx installed, grab it from - echo.http://sphinx-doc.org/ - exit /b 1 + echo. + echo.The 'sphinx-build' command was not found. Make sure you have Sphinx + echo.installed, then set the SPHINXBUILD environment variable to point + echo.to the full path of the 'sphinx-build' executable. Alternatively you + echo.may add the Sphinx directory to PATH. + echo. + echo.If you don't have Sphinx installed, grab it from + echo.http://sphinx-doc.org/ + exit /b 1 ) %SPHINXBUILD% -M %1 %SOURCEDIR% %BUILDDIR% %SPHINXOPTS% %O% diff --git a/src/tools/tc_utils/tc_diag/tc_diag.cc b/src/tools/tc_utils/tc_diag/tc_diag.cc index 1f8e483cf7..9cd5036b3f 100644 --- a/src/tools/tc_utils/tc_diag/tc_diag.cc +++ b/src/tools/tc_utils/tc_diag/tc_diag.cc @@ -149,7 +149,7 @@ __attribute__((noreturn)) static void usage(int exit_code) { << ") ***\n\n" << "Usage: " << program_name << "\n" << "\t-data domain tech_id_list [ file_1 ... file_n | file_list ]\n" - << "\t-deck file\n" + << "\t-deck path\n" << "\t-config file\n" << "\t[-outdir path]\n" << "\t[-log file]\n" @@ -162,7 +162,7 @@ __attribute__((noreturn)) static void usage(int exit_code) { << "\t\t\ta list of files to be used.\n" << "\t\t\tSpecify \"-data\" once for each data source (required).\n" - << "\t\t\"-deck source\" is the ATCF format data source " + << "\t\t\"-deck path\" is the ATCF format data source " << "(required).\n" << "\t\t\"-config file\" is a TCDiagConfig file to be used " diff --git a/src/tools/tc_utils/tc_gen/tc_gen.cc b/src/tools/tc_utils/tc_gen/tc_gen.cc index d9800ed0c9..8263dda38d 100644 --- a/src/tools/tc_utils/tc_gen/tc_gen.cc +++ b/src/tools/tc_utils/tc_gen/tc_gen.cc @@ -2575,31 +2575,31 @@ __attribute__((noreturn)) static void usage(int exit_code) { << ") ***\n\n" << "Usage: " << program_name << "\n" - << "\t-genesis source\n" - << "\t-edeck source\n" - << "\t-shape source\n" - << "\t-track source\n" + << "\t-genesis path\n" + << "\t-edeck path\n" + << "\t-shape path\n" + << "\t-track path\n" << "\t-config file\n" << "\t[-out base]\n" << "\t[-log file]\n" << "\t[-v level]\n\n" - << "\twhere\t\"-genesis source\" is one or more ATCF genesis " + << "\twhere\t\"-genesis path\" is one or more ATCF genesis " << "files, an ASCII file list containing them, or a top-level " << "directory with files matching the regular expression \"" << atcf_gen_reg_exp << "\" (required if no -edeck or -shape).\n" - << "\t\t\"-edeck source\" is one or more ensemble model output " + << "\t\t\"-edeck path\" is one or more ensemble model output " << "files, an ASCII file list containing them, or a top-level " << "directory with files matching the regular expression \"" << atcf_reg_exp << "\" (required if no -genesis or -shape).\n" - << "\t\t\"-shape source\" is one or more genesis area shapefiles, " + << "\t\t\"-shape path\" is one or more genesis area shapefiles, " << "an ASCII file list containing them, or a top-level " << "directory with files matching the regular expression \"" << gen_shp_reg_exp << "\" (required if no -genesis or -edeck).\n" - << "\t\t\"-track source\" is one or more ATCF track " + << "\t\t\"-track path\" is one or more ATCF track " << "files, an ASCII file list containing them, or a top-level " << "directory with files matching the regular expression \"" << atcf_reg_exp << "\" for the verifying BEST and operational " diff --git a/src/tools/tc_utils/tc_rmw/tc_rmw.cc b/src/tools/tc_utils/tc_rmw/tc_rmw.cc index 458e07c311..351d576d5f 100644 --- a/src/tools/tc_utils/tc_rmw/tc_rmw.cc +++ b/src/tools/tc_utils/tc_rmw/tc_rmw.cc @@ -119,7 +119,7 @@ __attribute__((noreturn)) static void usage(int exit_code) { << ") ***\n\n" << "Usage: " << program_name << "\n" << "\t-data file_1 ... file_n | file_list\n" - << "\t-deck file\n" + << "\t-deck path\n" << "\t-config file\n" << "\t-out file\n" << "\t[-log file]\n" @@ -129,10 +129,10 @@ __attribute__((noreturn)) static void usage(int exit_code) { << "specifies the gridded data files or an ASCII file " << "containing a list of files to be used (required).\n" - << "\t\t\"-deck source\" is the ATCF format data source " + << "\t\t\"-deck path\" is the ATCF format data source " << "(required).\n" - << "\t\t\"config_file\" is a TCRMWConfig file to be used " + << "\t\t\"-config file\" is a TCRMWConfig file to be used " << "(required).\n" << "\t\t\"-out file\" is the NetCDF output file to be written "