ctdam package

Subpackages

Submodules

Module contents

class ctdam.Casts(path_to_data='', ctd_data=[], processing_info={}, pattern='', plot=False, show_plot=True, plot_dir='htmls', use_multiprocessing=True)[source]

Bases: UserList

Work on multiple CTD casts simultaneously.

Parses all known CTD data formats if a file path given. Can also be initialized with instances of the cf-compliant xarray dataset structure used throughout ctdam. After parsing, straightforward processing and plotting can be run on all data. In the end, a grand data table of all concatenated casts can be written to disk as .tsv file. The cpu-heavy actions (converting and processing) can be executed in parallel, using multithreading.

Parameters:
  • path_to_data (Path | str) – Path to target files

  • ctd_data (list[Dataset]) – A list of target xarray Dataset objects

  • processing_info (dict) – Processing configuration

  • pattern (str) – A file pattern to filter the target files with

  • plot (bool) – Whether to create .html plots of the target files

  • show_plot (bool) – Whether to display the plots in a browser

  • plot_dir (Path | str) – The directory to store the plots in

  • use_multiprocessing (bool) – Whether to parallelize conversion and processing

data[source]

A list of CTDData objects

Type:

list

processing_info[source]

A processing workflow as used in Procedure

Type:

dict

path_to_data[source]

The input path to the data file(s)

Type:

Path

anomalous_data[source]

A list of data files that seem to be suspiciously small

Type:

list

cruise[source]

The name of the cruise the data files were collected in

Type:

str

sensor_info[source]

The sensor structure as generated by get_unique_sensor_data()

Type:

list

plot_dir[source]

A file path where generated data plots are saved to

Type:

Path

add_cast_numbers()[source]

Numbers the data files inside this instance

get_cruise_bottles(path_to_bl_files='', path_to_write='')[source]

Returns a DataFrame with all bottles closed during a single cruise.

Parameters:
  • path_to_bl_files (Path | str) – A file path to .bl files, if not in the same directory

  • path_to_write (Path | str) – File path to an output .csv file

Return type:

DataFrame

convert(file)[source]

Converts known CTD data files.

Can work with hex2py processing settings and hides all warnings.

Parameters:

file (Path) – A file path to a convertable CTD data file

Return type:

A parsed xarray Dataset

check_converted_data()[source]

Basic output data size check.

Catches very obvious test or wrong CTD data. These are saved in an anomalies attribute to allow human intervention.

Return type:

list[Dataset]

read_sensor_info()[source]

Parses all sensor metadata in one structure.

Usually, the sensor layout does not change during one cruise. But if it does, this function should detect that change and document it with a new list that displays the cast number and the differing sensor metadata.

process(processing_info, target_files=[])[source]

Applies the given processing workflow to all CTD data.

Uses multiprocessing for parallel processing and tqdm to display the progress.

Parameters:
  • processing_info (dict) – Processing parameters

  • target_files (list[Dataset]) – The input CTD data to process

Return type:

list[Dataset | None]

plot(show_plot=True)[source]

Creates .html plots of the target data.

Uses bokeh for interactive plotting and creates one ‘main’ .html that incorporates all individual cast plots.

Parameters:

show_plot (bool) – Whether to display the plots directly (Default value = True)

to_tsv(file_name=None)[source]

Exports the target file data into one great .tsv file.

Parameters:

file_name (str | Path | None) – The file name to write the tsv to (Default value = None)

ctdam.get_cast_borders(pressure, downcast_only=True, min_size_factor=0.01, min_soak_window=100, max_fd_quotient=6, prominence_divisor=7, win_size_divisor=500, min_velocity_quotient=15, min_velocity=0.045)[source]

Calculates start and end points of one CTD cast.

Uses first (fd) and second derivatives (sd) for that. Relies on carefully fine-tuned parameters that are set as default values. These can be fit to any kind of CTD data.

Parameters:
  • pressure (ndarray) – Pressure array

  • downcast_only (bool) – Whether to only work with downcast data (Default value = True)

  • min_size_factor (float) – Factor to check final dataset size against (Default value = 0.01)

  • min_soak_window (int) – Downcast_start: minimum size of soaking window (Default value = 100)

  • max_fd_quotient (int) – Downcast_start: Cut-off of fd height (Default value = 6)

  • prominence_divisor (int) – Downcast_start: Minimum size of sd peak prominence (Default value = 7)

  • win_size_divisor (int) – Downcast_start: Search window size to check fd means (Default value = 500)

  • min_velocity_quotient (int) – Downcast_start: Minimum velocity cut-off (Default value = 15)

  • min_velocity (float) – Downcast_start: Minimum velocity cut-off (Default value = 0.045)

Return type:

dict

ctdam.parse(file_path, downcast_only=False)[source]

Parse different file types to a cf-compliant xarray Dataset.

Can handle Seabirds .cnv and .hex file formats and Sea&Suns .TOB file format.

Parameters:

file_path (Path | str) – The path to the ctd data file

Return type:

Dataset

ctdam.plot(input, print_plot=True, output_directory='html', output_name='', html_title='', overwrite=False, no_new_plots=False, size_limit=10, filter='', show_html=True, config_path='vis_config.toml', file_type='cnv', use_multiprocessing=True, **kwargs)[source]

Display CTD data as interactive bokeh plots inside the web browser.

Single files or datasets will result in simnple plots. Multiple ones will all be individually plotted and a main entry html file will point to these individual plot files.

Parameters:
  • input (Path | str | Dataset | list) – The directory to look for data files to plot

  • print_plot (bool) – Whether to write plot files to disk (Default value = True)

  • output_directory (Path | str) – The directory to save .html file to (Default value = “html”)

  • output_name (str) – The name of the main html file (Default value = “main.html”)

  • html_title (str) – The header of the main html (Default value = “”)

  • overwrite (bool) – Whether to overwrite existing html plot files (Default value = False)

  • no_new_plots (bool) – Whether to not overwrite existing plot htmls (Default value = False)

  • size_limit (int) – Data file size limit in MB (Default value = 10)

  • filter (str) – A search filter for files (Default value = “”)

  • show_html (bool) – Whether to open main html in browser (Default value = True)

  • config_path (Path | str) – The path to vis configuration info (Default value = “vis_config.toml”)

  • file_type (str) – The file type to search for (Default value = “cnv”)

  • use_multiprocessing (bool) – Whether to use paralleliztion for plotting (Default value = True)

  • kwargs – All additional parameters will be given to basic_bokeh_plot

ctdam.process(input='', modules=['loopremoval', 'wildedit', 'wfilter', 'alignctd', 'celltm', 'binavg'], other_settings={}, use_multiprocessing=True, **kwargs)[source]

Run processing workflows on CTD data in file or memory form.

Does support multiple ‘modes’, depending on the input. 1) given a path to a directory, all parse-able CTD data formats found inside that directory will be parsed and processed, according to the workflow settings. 2) a path to a file will lead to the parsing and processing of that file. 3) similarly, the input could also consist of already parsed data, as single xarray dataset or a list of datasets.

The output will be processed cf-compliant xarray datasets.

Parameters:
  • input (Path | str | Dataset | list) – The data source to process

  • modules (dict | list) – The processing modules to apply to the data

  • other_settings (dict) – Processing configuration to use

  • use_multiprocessing (bool) – Whether to parallelize the operations

  • kwargs – Will be parsed as additional configuration options

Return type:

Dataset | List[Dataset]