- Updated the file schemas (PR #94):
src/api: The common dataset is now split into two censored datasets:censored_split1andcensored_split2. The method component should now be run twice, once for each split. The control methods return both splits at once. The metrics now receive both splits as input.src/control_methods: Control methods now need to return bothintegrated_split1andintegrated_split2.src/metrics: Metrics now receive bothintegrated_split1andintegrated_split2as input.src/workflows: Update the workflows accordingly.
-
Initialised repository with API files, readme, and a few test resources.
-
Added
methods/harmonypycomponent (PR #3). -
Added
methods/limma_remove_batch_effectcomponent (PR #7). -
Added three negative control methods (PR #8):
control_methods/shuffle_integrationcontrol_methods/shuffle_integration_by_batchcontrol_methods/shuffle_integration_by_cell_type
-
Added
metrics/emd_per_samplescomponent (PR #9). -
Added
methods/combatcomponent (PR #25). -
Added
control_methods/no_integration(PR #26). -
Added
metrics/n_inconsistent_peaks+ some new helper functions (PR #31). -
Added
methods/cycombine_nocontrol(PR #35). -
Added
metrics/average_batch_r2+ helper function (PR #36). -
Added
methods/gaussnorm(PR #45). -
Added
methods/cytornom_controls(PR #50). -
Added processing scripts for datasets (PR #55).
-
Added
metrics/flowsom_mapping_similarity(PR #59). -
Added
methods/mnn(PR #75). -
Updated cyCombine (PR #78):
- Added cyCombine with all controls (
methods/cycombine_all_controls) - Added cyCombine with one control - samples from only one condition (
methods/cycombine_one_control). - Added parameters to tune cyCombine.
- Added cyCombine with all controls (
-
Added
metrics/cms(PR #79). -
Added
methods/batchadjust_one_controlandmethods/batchadjust_all_controls(PR #82). -
Updated CytoNorm methods (PR #84):
- Added CytoNorm with one control - samples from only one condition (
methods/cytonorm_one_control) - Added CytoNorm with aggregate of samples as controls (
methods/cytonorm_no_controls). - Added parameters to tune CytoNorm.
- Added CytoNorm with one control - samples from only one condition (
-
Added CytoNorm correction to a goal batch (PR #92).
-
Added cyCombine correction to a reference batch (PR #90).
-
Added
metrics/bras(PR #91). -
Added Seurat rPCA (PR #95).
-
Added processing scripts for CLL dataset (PR #106).
-
Added new metric
ratio_inconsistent_peaks(PR #114). -
Added processing scripts for Lille dataset and remove ones for CLL dataset (PR #118).
-
Added config and run scripts for running the benchmark on WEHI HPC (PR #119).
-
Added utility scripts to pull intermediate files (PR #119).
-
Added
scripts/fetch_intermediate_files.py, a single script that consolidates the previous two-step process of parsing a Nextflow log and copying out intermediate files (PR #126):- Auto-detects whether the log is from a SLURM or AWS Batch run.
- For SLURM runs, copies
.h5adfiles directly from the local work directory. - For AWS Batch runs, downloads files via
common/scripts/fetch_task_runand skips files that are already present. - Organises output under
<dataset>/method_out/and<dataset>/metric_out/. - Accepts the log file as a positional argument; output directory and optional CSV dump are configurable via flags.
-
Added
metrics/functional_marker_preservationcomponent (PR #126):functional_marker_preservation_wilcoxon: proportion of biologically significant functional marker group differences (e.g. WT vs KO) preserved after batch integration. A two-sided Wilcoxon rank-sum test compares the two biological groups on per-sample mean expression for each (marker, cell type) pair. A pair enters the unintegrated baseline if significant (p ≤ 0.1) in both batches, and is counted as preserved if significant in both integrated splits.functional_marker_preservation_cohens_d: mean absolute change in Cohen's d effect size for every (marker, cell type) pair significant in the unintegrated baseline. Cohen's d is computed per batch on the unintegrated data and averaged, then computed separately per integrated split; the metric is the mean absolute deviation of each split's Cohen's d from the unintegrated ground truth, so drift is tracked even for pairs that lost significance after integration.
-
Added
scripts/adhoc_runs/run_metric_adhoc.pyandrun_metric_adhoc.shfor re-running a single metric against existing pipeline output during development, without re-running the full Nextflow pipeline. Patches the metric script's VIASH block with the given input/output paths and runs it as a subprocess; discovers(dataset, method)pairs from<input-dir>/<dataset>/method_out/<method>_split1.h5adand supports skipping datasets/methods already processed (PR #126).
-
Removed
methods/cytovifrom the benchmark. The implementation is preserved in theadd-cytovi-implementationbranch to be revisited in the near future (PR #124). -
Updated
metrics/lisi(PR #126):- iLISI is still computed globally across all cells per split (not per group), but is now skipped and returns NaN for a split if batch is confounded by group in any group (i.e. a group whose cells all come from a single batch, so cross-batch mixing can't be assessed).
- Moved the iLISI/cLISI computation out of
script.pyinto a newhelper.py(compute_ilisi,compute_clisi,_check_batch_group_confounding), and expanded the outputunsto include per-cell iLISI/cLISI values and cell IDs for each split.
-
Updated
methods/harmonypyto use the PyTorch GPU backend (harmonypy>=0.2.0) (PR #126):- Switched Docker image to
openproblems/base_pytorch_nvidia:1.0.0. - Added
device=Nonetorun_harmonyfor automatic GPU detection (CUDA -> MPS -> CPU). - Removed epsilon workaround as numerical stability is now handled internally by harmonypy.
- Updated Nextflow label to
lowcpu, gpu.
- Switched Docker image to
-
Updated file schema (PR #18):
- Add is_control obs to indicate whether a cell should be used as control when correcting batch effect.
- Removed donor_id obs from unintegrated censored.
- Removed to_correct var from everything except common_dataset. All datasets now will only contain markers that need to be corrected.
-
Reupdated the file schema (PR #19):
- Included changes in PR #21: data Processor component partitions cells between unintegrated(censored) and validation.
- Add back to_correct var to every file except integrated to reflect the real world batch correction workflow better.
- Reverted PR #18 to retain only the 1st two changes (add is_control and remove donor_id from unintegrated_censored).
-
Changed emd to emd_mean and emd_max (PR #27).
-
Updated output anndata for methods to return all vars - corrected or not (PR #28).
-
Added positive control (PR #30).
-
Rewrote EMD metric so it no longer relies on implementation in cytonormpy (PR #48):
- EMD is now calculated for all cell types as well.
- Returned values are average and max across all cell types (exclude the one calculated agnostic of cell types), average and max across all donors (mean and max of values computed agnostic of cell types).
-
Bump Viash version to 0.9.4 (PR #61, PR #62).
-
Added EMD vertical global metric and split perfect integration into horizontal and vertical for computing horizontal and vertical metrics (PR #63).
-
Fix problems identified during a full run (PR #99).
-
Update CytoVI (PR #114).
-
Update CytoVI to normalise using minmax scaler fitted on batch 1 post correction (PR #119).
-
Update batchadjust, cytonorm to use HPC temp dir if the environment variable is set or else default to what is set by viash. This is to prevent collision in temp files when the jobs are running (PR #119).
-
Update ratio inconsistent peaks to handle edge cases where methods return only zero values for a marker/cell type/donor combination, causing sd to be zero and division by zero (PR #119).
-
One control and no control method will only get either samples from one control plus non-control samples or just no control samples. They will no longer be given access to other samples to correct or to train the model. Notably, the included control samples may still be corrected (PR #119).
-
Change temp folder for methods which rely on writing out FCS files. Temp folders are now created by a new helper function which will create a subdirectory under
meta[["temp_dir"]]. This will be used as the temp directory (PR #119). -
Change inconsistent peaks metrics to consistent peaks (PR #119).
-
Enabled unit tests (PR #2).
-
Added integrated test resource (PR #5).
-
Updated file description in yaml file (PR #15).
-
Updated scripts to enable running benchmark on seqera (PR #29).
-
Updated project description (PR #42).
-
Implemented function for writing FCS files from an anndata object (PR #49).
-
Changed cytonorm and cycombine clustering to use lineage markers (PR #54).
-
The metric
flowsom_mapping_similaritynow works at the cluster level (PR #68). -
Updated vertical EMD (PR #70):
- Metric is computed in group (condition) specific manner.
- Split the metric into global and per cell type.
- Refactored horizontal EMD.
-
Updated
metrics/flowsom_mapping_similarityto compute mapping similarity bidirectionally (split1→split2 and split2→split1) and to export per-donor cluster×cell_type absolute-difference matrices in the output AnnDatauns(PR #126). -
Schema defined in
src\apihas been modified to include dataset specific parameters (PR #71). -
Disabled parallelization of
metrics/cms+ cms distributions are now kept in the output AnnData object (PR #83). -
Add arguments for including/excluding methods and metrics in the benchmarking workflow (PR #100).
-
Removed EMD max from calculation (PR #113).
-
Tune the resource requirement for each method (PR #119).
- Low time, mem, cpu for control methods.
- Mid time, mem, cpu for most methods, except below.
- High (or very high) time, mem, cpu for computationally expensive methods like rPCA.
-
Change n_inconsistent_peaks output to float and add R2 to main.nf (PR #40).
-
Changed naming 'gaussNorm' to 'gaussnorm' (PR #47).
-
Fixed cyCombine so it now batch correct unnormalised data rather than normalised data (PR #58).
-
Fixed FlowSOM mapping similarity metric (PR #64).
-
Fixed get_obs_var_for_integrated function in helper.R giving out error in mac (PR #65).
-
Remove unlabelled cells when computing n_inconsistent_peaks metric (PR #69).
-
Fix vertical EMD (PR #76):
- Return NA if there are less than 2 samples per group in the data.
- Refactoring and introduced "global" variable for output.
-
Add MNN to run script (PR #77).
-
Added
argument_groupsfield inmethods/mnnconfig file (PR #81). -
Fix unparseable characters in EMD metrics (PR #87).
-
Fix missing anndata in yaml file and set the base_r docker image version to 1 instead of 1.0.0 (PR #89).
-
Fix bug in control methods (PR #107, #108, #109).
- All control methods are updated to cater the new schema.
- All control methods are re-enabled. Selectively disable them when running the pipeline using method exclude.
-
Fix bug in EMD where nan cannot be written out and added sklearn dependency for cytovi (PR #110).
-
Fix bug in EMD vertical where sample combination was malformed (PR #113)
-
Fix lisi inconsistent naming (PR #117) for issue #116.
-
Fix bug in perfect integration where if batch is str (not int), it only returns control samples (PR #119).
-
Fix bug in batchadjust needing "Batch_" in the sample names for non-control samples (PR #119).
-
Fix bug in cytonorm to mid where recompute was set to FALSE. It is now set to TRUE (PR #119).
-
Remove transpose in harmonypy as new updates to harmonypy no longer need the transpose (PR #119).
-
Fix bug in get_obs_var_for_integrated to handle the cases where batch column in obs is str and thus can't be directly overriden (new values given by get_donor_batch_map is int) (PR #119).
-
Update flowsom mapping similarity so we subset to just markers to correct, and lisi to remove control samples and unlabelled cells (PR #119).
-
Point
scripts/run_benchmark/wehi_hpc/run_full_hpc.shatbuild/maininstead ofbuild/update_ilisi, and label the seqera full run asfullinstead oftest_subset(PR #129). -
Fix
average_batch_r2andflowsom_mapping_similaritywritingmetric_idsandmetric_valuesas scalars instead of lists, as required byfile_score.yaml(PR #130). -
Clean up stale mock parameters and dead code (PR #133).
-
Fix bug in
average_batch_r2where the R2 was computed on all cell types of a donor at once instead of on each cell type separately (PR #127).