1 Introduction

LIPIDIFy provides an end-to-end toolkit for mass-spectrometry lipidomics data analysis: missing-value imputation, batch-effect correction, multiple normalization strategies, differential abundance analysis (limma and edgeR), gene-set enrichment analysis, and extensive visualization. Lipid names are automatically classified by class, subclass, and fatty-acid saturation, and the package supports user-defined classification schemes and enrichment sets for lipidomes that don’t fit a fixed ontology.

1.1 Why LIPIDIFy?

Several excellent packages already support lipidomics analysis in Bioconductor, most notably lipidr, which centres on a SummarizedExperiment-derived class for annotation, normalization, and class-level statistical testing. LIPIDIFy is a complementary tool rather than a replacement, aimed at two audiences at once:

  • Bench biologists, via launch_lipidomics_app(), an interactive Shiny dashboard that requires no R programming: upload a CSV, click through normalization/QC/differential/enrichment steps, and export a report.
  • Bioinformaticians, via the same underlying functions used programmatically, for reproducible, scriptable analyses.

Relative to existing workflows, LIPIDIFy additionally emphasises:

  • Flexible classification and enrichment. Lipid classes are inferred automatically from standard nomenclature, but users can supply their own classification table (any hierarchy, any number of columns) and custom enrichment sets, rather than being tied to one fixed ontology.
  • Side-by-side pipeline comparison. create_pipeline_plot() lets users visually compare two full normalization pipelines (e.g. TIC+Log2 vs. PQN+Log2) before committing to one, which is otherwise a manual, repetitive step.
  • Imputation and batch correction as first-class steps, with multiple interchangeable methods (impute_missing_values(), correct_batch_effects()) integrated into the same workflow as normalization and differential analysis, rather than requiring separate tooling.

1.2 Interoperability with the Bioconductor ecosystem

Internally, LIPIDIFy represents data as plain matrices and data frames (list(data, metadata, numeric_data)) rather than a custom S4 class, which keeps the API approachable for users who are not already familiar with Bioconductor’s class system. This representation converts trivially to a SummarizedExperiment, so results can be handed off to other Bioconductor tools (or a lipidr LipidomicsExperiment, which extends SummarizedExperiment) for further analysis. See the Interoperability section near the end of this vignette for a worked example.

1.3 Installation

if (!requireNamespace("BiocManager", quietly = TRUE)) {
  install.packages("BiocManager")
}

BiocManager::install("LIPIDIFy")
library(LIPIDIFy)

The same workflow below is available without any coding via the Shiny application:

launch_lipidomics_app()

2 Example Data

This vignette uses a real, publicly available lipidomics dataset rather than simulated data: Metabolomics Workbench study ST001359, “Monophasic lipidomics extraction in cancer cell lines” (Rodriguez Blanco G., Beatson Institute for Cancer Research; Project PR000929, Analysis AN002263; licensed CC BY 4.0). It profiles HepG2 hepatocellular carcinoma cells by LC-MS (reversed phase, positive ion mode) under two conditions: untreated controls and cells treated with an SDC (syndecan) inhibitor, 3 replicates each. The processed data matrix is bundled with the package at inst/extdata/ST001359_lipidomics.csv; the script that downloaded and prepared it is inst/scripts/prepare_ST001359_lipidomics.R.

file_path <- system.file("extdata", "ST001359_lipidomics.csv", package = "LIPIDIFy")

raw_data <- load_lipidomics_data(
  file_path,
  metadata_columns = c("Sample Name", "Sample Group")
)

cat("Samples  :", nrow(raw_data$numeric_data), "\n")
#> Samples  : 6
cat("Lipids   :", ncol(raw_data$numeric_data), "\n")
#> Lipids   : 446
cat("Groups   :", paste(unique(raw_data$metadata$`Sample Group`), collapse = ", "), "\n")
#> Groups   : Control, SDC_i

load_lipidomics_data() returns a named list with three elements:

  • $numeric_data - matrix of lipid abundances (samples x lipids)
  • $metadata - data frame of sample metadata
  • $data - the original full data frame
head(raw_data$metadata, 6)
#>                       Sample Name Sample Group
#> VV_13_HEpG2_C1     VV_13_HEpG2_C1      Control
#> VV_14_HEpG2_C2     VV_14_HEpG2_C2      Control
#> VV_15_HEpG2_C3     VV_15_HEpG2_C3      Control
#> VV_16_HEpG2_SDC1 VV_16_HEpG2_SDC1        SDC_i
#> VV_17_HEpG2_SDC2 VV_17_HEpG2_SDC2        SDC_i
#> VV_18_HEpG2_SDC3 VV_18_HEpG2_SDC3        SDC_i
raw_data$numeric_data[, 1:4]
#>                  AC(12:0) AC(14:0) AC(14:1) AC(16:0)
#> VV_13_HEpG2_C1       7.82     9.87     8.71    10.59
#> VV_14_HEpG2_C2       7.83     9.95     8.88    10.70
#> VV_15_HEpG2_C3       7.88     9.98     8.83    10.79
#> VV_16_HEpG2_SDC1     7.31     9.61     6.60    10.83
#> VV_17_HEpG2_SDC2     7.62     9.65     6.89    10.75
#> VV_18_HEpG2_SDC3     7.83    10.19     7.33    11.19

Values are the study’s own “peak area, normalized” units, which the submitters already log2-scaled; we use them unchanged (not re-normalized or back-transformed) to preserve fidelity to the deposited data. This matters below: a second application of Log2 or TIC normalization would not be meaningful on data that is already log-scaled and normalized, so we apply only light median-centring further on.

If you would rather experiment with a larger, fully synthetic dataset (20 samples, 4 groups, un-transformed intensities), generate_example_data() creates one on the fly – it is used in this vignette only for the normalization-pipeline comparison in the next section, where raw-scale intensities make the effect of each method clearer.


3 Lipid Classification

classify_lipids() parses lipid names and assigns three classification levels:

  • LipidGroup - broad category (e.g., Glycerophospholipids, Sphingolipids)
  • LipidType - specific subclass (e.g., Phosphatidylcholine)
  • Saturation - SFA (0 double bonds), MUFA (1), PUFA (> 1)
lipid_names <- colnames(raw_data$numeric_data)
classification <- classify_lipids(lipid_names)

head(classification, 8)
#>                          Lipid         LipidGroup LipidType Saturation
#> 1                     AC(12:0) Other/Unclassified     Other        SFA
#> 2                     AC(14:0) Other/Unclassified     Other        SFA
#> 3                     AC(14:1) Other/Unclassified     Other       MUFA
#> 4                     AC(16:0) Other/Unclassified     Other        SFA
#> 5                     AC(18:0) Other/Unclassified     Other        SFA
#> 6                     AC(18:1) Other/Unclassified     Other       MUFA
#> 7 Alkenyl-TG(P-18:0_16:0_18:1) Other/Unclassified     Other       MUFA
#> 8           Alkenyl-TG(P-48:0) Other/Unclassified     Other        SFA
cat("Lipid groups:\n")
#> Lipid groups:
print(sort(table(classification$LipidGroup), decreasing = TRUE))
#> 
#> Glycerophospholipids   Other/Unclassified        Glycerolipids 
#>                  179                  114                   79 
#>        Sphingolipids        Sterol Lipids 
#>                   60                   14

cat("\nFatty-acid saturation:\n")
#> 
#> Fatty-acid saturation:
print(table(classification$Saturation))
#> 
#> MUFA PUFA  SFA 
#>   85  284   77

Real-world lipid identifiers are rarely as tidy as textbook nomenclature. Here, roughly a quarter of features (e.g. BMP, CL, acylcarnitines, ether lipids reported as Plasmanyl-/Plasmenyl-) fall outside the built-in pattern set and are labelled "Other/Unclassified". Rather than guessing, you can supply your own classification table with load_custom_classification() (any number of columns, any hierarchy) to cover exactly the lipid space of your dataset.


4 Raw Data Visualization

Examining distributions before normalization helps detect outlier samples, batch effects, and scaling differences.

raw_list <- list(numeric_data = raw_data$numeric_data, metadata = raw_data$metadata)

visualize_raw_data_improved(
  raw_list,
  plot_type    = "boxplot",
  view_mode    = "sample",
  metadata     = raw_data$metadata,
  group_column = "Sample Group"
)
Per-sample intensity distributions before normalization. Boxes are coloured by group.

Figure 1: Per-sample intensity distributions before normalization
Boxes are coloured by group.

visualize_raw_data_improved(
  raw_list,
  plot_type    = "density",
  view_mode    = "sample",
  metadata     = raw_data$metadata,
  group_column = "Sample Group"
)
Intensity density curves before normalization, one curve per sample.

Figure 2: Intensity density curves before normalization, one curve per sample


5 Normalization

5.1 What methods are available?

get_normalization_methods()
#>  [1] "TIC"        "PQN"        "Quantile"   "VSN"        "Log2Median"
#>  [6] "Median"     "Mean"       "Log2"       "Log10"      "Sqrt"      
#> [11] "None"

5.2 VSN and Log2Median are different methods

Two of these deserve a note, because their names invite confusion:

  • "Log2Median" applies a fixed transformation: log2(x + 1), then a per-sample median shift onto a common baseline, back-transformed to the original scale. Nothing is estimated from the data.
  • "VSN" performs genuine variance-stabilising calibration with vsn::justvsn() from the Bioconductor vsn package (Huber et al., Bioinformatics 18, S96–S104, 2002). It fits an arcsinh (generalised-log) transformation whose parameters are estimated by robust maximum likelihood, simultaneously calibrating between-sample differences and stabilising the variance-versus-mean relationship.

VSN may be appropriate for positive, intensity-like measurements whose variance grows with the mean – a common pattern in raw MS lipidomics data. It is not simply another name for a log transformation, and it is not automatically the best choice for every lipidomics dataset; compare pipelines on your own data before committing to one.

Practical restrictions:

  • The Bioconductor vsn package must be installed (BiocManager::install("vsn")). If it is missing, "VSN" raises an error – it never silently falls back to another method.
  • At least 2 samples and at least 42 lipid features are required (vsn’s minDataPointsPerStratum default), and values must be finite. Infinite values, samples or features that are entirely missing, and samples with fewer than two observed values are rejected with a descriptive error.
  • Missing values are permitted: vsn fits on the observed values and returns NA in the same positions. They are never imputed.
  • Negative values are passed to vsn unchanged but produce a warning, since they usually indicate data that has already been log-transformed or background-subtracted.
  • The output is already on a log-like scale, so do not chain a "Log2" step after "VSN".
synthetic_obj_vsn <- load_lipidomics_data_from_df(generate_example_data())

vsn_norm <- apply_normalizations(synthetic_obj_vsn$numeric_data, "VSN")
dim(vsn_norm) # samples x lipids - orientation and dimnames preserved
#> [1]  20 301

# Distinct from Log2Median, which is a different method:
l2m_norm <- apply_normalizations(synthetic_obj_vsn$numeric_data, "Log2Median")
isTRUE(all.equal(vsn_norm, l2m_norm))
#> [1] FALSE

5.3 Comparing pipelines side-by-side

It is good practice to compare normalization strategies visually before committing to one. This is most informative on raw-scale intensities, so we illustrate it here with the built-in synthetic dataset, comparing TIC+Log2 (a common starting point) against PQN+Log2 (more robust when large fold-changes are expected):

synthetic_df <- generate_example_data()
synthetic_obj <- load_lipidomics_data_from_df(synthetic_df)

pipeline1 <- apply_normalizations(synthetic_obj$numeric_data, c("TIC", "Log2"))
pipeline2 <- apply_normalizations(synthetic_obj$numeric_data, c("PQN", "Log2"))

p1 <- create_pipeline_plot(pipeline1,
  title     = "Pipeline 1: TIC + Log2",
  metadata  = synthetic_obj$metadata,
  plot_type = "boxplot"
)

p2 <- create_pipeline_plot(pipeline2,
  title     = "Pipeline 2: PQN + Log2",
  metadata  = synthetic_obj$metadata,
  plot_type = "boxplot"
)

gridExtra::grid.arrange(p1, p2, ncol = 1)
Side-by-side comparison of two normalization pipelines applied to the same raw, synthetic data.

Figure 3: Side-by-side comparison of two normalization pipelines applied to the same raw, synthetic data

5.4 Applying normalization to the real dataset

The ST001359 data is already log2-scaled and normalized by the original submitters, so we apply only median-centring here rather than a second TIC/PQN/Log2 pass, which would not be meaningful on top of an existing log-scale normalization:

norm_data <- apply_normalizations(raw_data$numeric_data, methods = "Median")

dim(norm_data) # samples x lipids - same dimensions as raw data
#> [1]   6 446
norm_list <- list(numeric_data = norm_data, metadata = raw_data$metadata)

visualize_raw_data_improved(
  norm_list,
  plot_type    = "boxplot",
  view_mode    = "sample",
  metadata     = raw_data$metadata,
  group_column = "Sample Group"
)
Per-sample distributions after median-centring.

Figure 4: Per-sample distributions after median-centring


6 Missing Value Imputation

Missing values in lipidomics data typically arise when a lipid’s signal falls below the instrument’s limit of detection. LIPIDIFy provides six imputation strategies via impute_missing_values(). Imputing after normalisation ensures that imputed values sit on the same scale as observed values.

cat("Missing values in this dataset:", sum(is.na(norm_data)), "\n")
#> Missing values in this dataset: 0
cat("Imputation methods available:\n")
#> Imputation methods available:
print(get_imputation_methods())
#> [1] "half_min" "min"      "zero"     "mean"     "median"   "knn"
# For demonstration: introduce 5% artificial missing values
demo_norm <- norm_data
set.seed(42)
demo_norm[sample(length(demo_norm), size = round(0.05 * length(demo_norm)))] <- NA
cat(
  "Missingness introduced:", sum(is.na(demo_norm)),
  sprintf("(%.1f%%)\n", 100 * mean(is.na(demo_norm)))
)
#> Missingness introduced: 134 (5.0%)

imputed_norm <- impute_missing_values(demo_norm, method = "half_min")
cat("Missing after imputation:", sum(is.na(imputed_norm)), "\n")
#> Missing after imputation: 0

Choosing a method:

Method Best for
half_min Most lipidomics data; models the LOD on normalised scale
knn Data missing at random (MAR); requires impute Bioconductor package
zero Only when absence of signal truly means zero
mean / median Small proportions of missing values

For the remainder of this vignette we use the complete normalised data without artificial missingness.


7 Batch Effect Correction

When samples are processed across multiple runs, instruments, or operators, technical batch effects can obscure biological differences. correct_batch_effects() removes these effects while preserving biological signal, and should be applied after normalisation and imputation, then verified with PCA. The ST001359 dataset has no recorded batch structure, so this step is illustrated but not evaluated here.

Two methods are available:

Method When to use
"limma" Default; fast and reliable for balanced designs
"combat" Better for large or unbalanced batch effects; requires sva package
# Simulate batch labels (real data would have this in the metadata file)
raw_data$metadata$Batch <- rep(c("Run1", "Run2"), each = 3)

norm_corrected <- correct_batch_effects(
  data_matrix  = norm_data,
  metadata     = raw_data$metadata,
  batch_column = "Batch",
  group_column = "Sample Group",
  method       = "limma"
)

# Verify: PCA should now show clustering by biology, not by batch
pca_corrected <- perform_pca(norm_corrected, raw_data$metadata, "Sample Group")
create_pca_plot_with_ellipses(
  pca_data           = pca_corrected$pca_data,
  variance_explained = pca_corrected$variance_explained,
  ellipse_type       = "confidence",
  title              = "PCA after Batch Correction"
)

8 Quality Control with PCA

PCA on normalized data should show samples clustering by biological group rather than by technical artefact.

pca_res <- perform_pca(norm_data, raw_data$metadata, "Sample Group")

create_pca_plot_with_ellipses(
  pca_data           = pca_res$pca_data,
  variance_explained = pca_res$variance_explained,
  ellipse_type       = "none",
  show_sample_labels = TRUE,
  title              = "PCA - Control vs. SDC Inhibition"
)
PCA of the normalized ST001359 data.

Figure 5: PCA of the normalized ST001359 data

With only 3 replicates per group, confidence ellipses are not meaningful here (ellipse_type = "none"); with larger sample sizes, ellipse_type = "confidence" draws 95% confidence regions per group.


9 Differential Analysis

9.1 Automatic pairwise contrasts

create_default_contrasts() generates all pairwise comparisons from the groups present in your data. You can also specify contrasts manually, or supply your own design matrix (see ?perform_differential_analysis) when the default ~0 + group design doesn’t fit your experiment (e.g. paired samples, additional covariates).

groups <- unique(raw_data$metadata$`Sample Group`)
contrasts <- create_default_contrasts(groups)
print(contrasts)
#> [1] "SDC_i - Control"

9.2 Choosing a method: limma or edgeR?

diff_res <- perform_differential_analysis(
  data_matrix    = norm_data,
  metadata       = raw_data$metadata,
  group_column   = "Sample Group",
  contrasts_list = contrasts,
  method         = "limma"
)

limma is the recommended default for mass-spectrometry lipidomics: it models continuous, approximately log-normal intensities directly.

edgeR is provided only as an exploratory alternative. edgeR’s negative-binomial model was designed for integer RNA-seq count data; applying it to continuous MS intensities is an approximation, and perform_differential_analysis(..., method = "edger") emits a warning to that effect every time it is called. Consider edgeR only if you specifically want to cross-check limma’s results against a second modelling assumption, or if you are integrating lipidomics into a broader pipeline that already standardises on edgeR for consistency – not as a first choice.

cat(sprintf("%-20s  %5s  %5s  %4s  %5s\n", "Contrast", "Total", "Sig", "Up", "Down"))
#> Contrast              Total    Sig    Up   Down
cat(strrep("-", 48), "\n")
#> ------------------------------------------------

for (nm in names(diff_res$results)) {
  res <- diff_res$results[[nm]]
  n_sig <- sum(res$adj.P.Val < 0.05, na.rm = TRUE)
  n_up <- sum(res$adj.P.Val < 0.05 & res$logFC > 0, na.rm = TRUE)
  n_down <- sum(res$adj.P.Val < 0.05 & res$logFC < 0, na.rm = TRUE)
  cat(sprintf("%-20s  %5d  %5d  %4d  %5d\n", nm, nrow(res), n_sig, n_up, n_down))
}
#> SDC_i - Control         446    283   146    137
first_contrast <- names(diff_res$results)[1]
res_df <- diff_res$results[[first_contrast]]
res_df <- res_df[order(res_df$adj.P.Val), ]

top10 <- utils::head(res_df[, c("logFC", "AveExpr", "adj.P.Val")], 10)
top10$logFC <- round(top10$logFC, 3)
top10$AveExpr <- round(top10$AveExpr, 2)
top10$adj.P.Val <- signif(top10$adj.P.Val, 3)

cat("Top 10 features in:", first_contrast, "\n\n")
#> Top 10 features in: SDC_i - Control
print(top10)
#>                 logFC AveExpr adj.P.Val
#> PC(37:0)        6.891    8.23  1.17e-07
#> PC(34:4)       -2.526   11.75  7.11e-06
#> PC(36:0)        2.530   12.46  7.11e-06
#> SHexCer(d38:5) -2.899    7.27  7.11e-06
#> DG(16:0_18:0)   3.992   11.85  9.85e-06
#> PG(16:0_18:0)   4.690    4.74  9.85e-06
#> PI(42:9)        1.890    8.26  9.85e-06
#> PC(34:0)        1.772   17.15  9.85e-06
#> LysoPC(17:0)    1.980    8.27  9.85e-06
#> BMP(38:6)       1.520    7.88  1.30e-05

10 Results Visualization

10.1 Volcano plot

Significant lipids (|logFC| > 1 and adj.P.Val < 0.05) are coloured by lipid class. The top 10 most significant are labelled.

create_volcano_plot_labeled(
  results             = diff_res$results[[first_contrast]],
  title               = paste("Volcano:", first_contrast),
  logfc_threshold     = 1,
  pval_threshold      = 0.05,
  top_labels          = 10,
  classification_data = classification,
  color_by            = "LipidGroup"
)
Volcano plot coloured by lipid class.

Figure 6: Volcano plot coloured by lipid class

10.2 Heatmap of significant features

sig_lipids <- rownames(res_df)[
  !is.na(res_df$adj.P.Val) &
    res_df$adj.P.Val < 0.05 &
    abs(res_df$logFC) > 1
]

if (length(sig_lipids) >= 5) {
  top_n <- min(30, length(sig_lipids))
  avail <- intersect(sig_lipids[seq_len(top_n)], colnames(norm_data))
  hm_mat <- t(norm_data[, avail, drop = FALSE])

  create_heatmap_robust(
    data_matrix  = hm_mat,
    metadata     = raw_data$metadata,
    group_column = "Sample Group",
    top_n        = top_n,
    title        = paste("Significant lipids:", first_contrast)
  )
} else {
  message("Fewer than 5 significant lipids at these thresholds.")
}

11 Lipid Expression

create_lipid_expression_barplot() shows per-sample abundance of individual lipids, sorted by group. It returns a single ggplot object when one lipid is selected, or a named list of ggplot objects when multiple lipids are selected.

top_lipids <- utils::head(rownames(res_df)[!is.na(res_df$adj.P.Val)], 3)
cat("Selected lipids:\n")
#> Selected lipids:
print(top_lipids)
#> [1] "PC(37:0)" "PC(34:4)" "PC(36:0)"

expr_result <- create_lipid_expression_barplot(
  data_matrix     = norm_data,
  metadata        = raw_data$metadata,
  selected_lipids = top_lipids,
  group_column    = "Sample Group",
  data_type       = "normalized"
)

if (inherits(expr_result, "ggplot")) {
  print(expr_result)
} else {
  for (lipid_plot in expr_result) print(lipid_plot)
}
Per-sample abundance of the 3 most significant lipids.

Figure 7: Per-sample abundance of the 3 most significant lipids

Per-sample abundance of the 3 most significant lipids.

Figure 8: Per-sample abundance of the 3 most significant lipids

Per-sample abundance of the 3 most significant lipids.

Figure 9: Per-sample abundance of the 3 most significant lipids


12 Enrichment Analysis

Enrichment analysis tests whether lipids in a particular class are collectively more increased or decreased than expected by chance, using the fgsea algorithm.

enrich_res <- perform_enrichment_analysis(
  results_list        = diff_res$results,
  classification_data = classification,
  min_set_size        = 3,
  max_set_size        = 500
)

cat("Categories per contrast:", paste(names(enrich_res[[1]]), collapse = ", "), "\n")
#> Categories per contrast: LipidGroup, LipidType, Saturation
grp_enrich <- enrich_res[[first_contrast]][["LipidGroup"]]

if (!is.null(grp_enrich) && nrow(grp_enrich) > 0) {
  grp_enrich <- grp_enrich[order(grp_enrich$pval), ]
  out <- utils::head(grp_enrich[, c("pathway", "NES", "pval", "padj", "size")], 5)
  out$NES <- round(out$NES, 3)
  out$pval <- signif(out$pval, 3)
  out$padj <- signif(out$padj, 3)
  print(out, row.names = FALSE)
} else {
  cat("No enrichment results available.\n")
}
#>               pathway    NES   pval   padj  size
#>                <char>  <num>  <num>  <num> <int>
#>         Glycerolipids -1.436 0.0194 0.0968    79
#>         Sphingolipids  1.043 0.3790 0.6260    60
#>         Sterol Lipids -0.957 0.5190 0.6260    14
#>    Other/Unclassified  0.930 0.5890 0.6260   114
#>  Glycerophospholipids  0.932 0.6260 0.6260   179
if (!is.null(grp_enrich) && nrow(grp_enrich) > 0) {
  create_enrichment_dotplot(
    enrichment_data = grp_enrich,
    title           = paste("Lipid Group Enrichment:", first_contrast),
    max_pathways    = 10
  )
}
Enrichment dot plot for the first contrast. Dot size reflects set size; colour reflects statistical significance.

Figure 10: Enrichment dot plot for the first contrast
Dot size reflects set size; colour reflects statistical significance.

Custom, user-defined pathway sets (not tied to LipidGroup/LipidType) can be supplied via load_custom_enrichment_sets() and passed as custom_sets.


13 Interoperability with SummarizedExperiment

LIPIDIFy’s list-of-matrices representation converts directly to a SummarizedExperiment, for interoperability with the rest of the Bioconductor ecosystem (or with lipidr, whose LipidomicsExperiment class extends SummarizedExperiment):

if (requireNamespace("SummarizedExperiment", quietly = TRUE)) {
  se <- SummarizedExperiment::SummarizedExperiment(
    assays  = list(normalized = t(norm_data)),
    colData = raw_data$metadata
  )
  print(se)
}
#> class: SummarizedExperiment 
#> dim: 446 6 
#> metadata(0):
#> assays(1): normalized
#> rownames(446): AC(12:0) AC(14:0) ... TG(62:3) TG(62:4)
#> rowData names(0):
#> colnames(6): VV_13_HEpG2_C1 VV_14_HEpG2_C2 ... VV_17_HEpG2_SDC2
#>   VV_18_HEpG2_SDC3
#> colData names(2): Sample Name Sample Group

From here, se can be passed to any Bioconductor function expecting a SummarizedExperiment.


14 Funding

This work was supported by a grant to A/Prof Karen Sheppard from the National Health and Medical Research Council of Australia (#2020050).


15 Session Information

sessionInfo()
#> R version 4.6.1 (2026-06-24)
#> Platform: x86_64-pc-linux-gnu
#> Running under: Ubuntu 24.04.5 LTS
#> 
#> Matrix products: default
#> BLAS:   /home/biocbuild/bbs-3.24-bioc/R/lib/libRblas.so 
#> LAPACK: /usr/lib/x86_64-linux-gnu/lapack/liblapack.so.3.12.0  LAPACK version 3.12.0
#> 
#> locale:
#>  [1] LC_CTYPE=en_US.UTF-8       LC_NUMERIC=C              
#>  [3] LC_TIME=en_GB              LC_COLLATE=C              
#>  [5] LC_MONETARY=en_US.UTF-8    LC_MESSAGES=en_US.UTF-8   
#>  [7] LC_PAPER=en_US.UTF-8       LC_NAME=C                 
#>  [9] LC_ADDRESS=C               LC_TELEPHONE=C            
#> [11] LC_MEASUREMENT=en_US.UTF-8 LC_IDENTIFICATION=C       
#> 
#> time zone: America/New_York
#> tzcode source: system (glibc)
#> 
#> attached base packages:
#> [1] stats     graphics  grDevices utils     datasets  methods   base     
#> 
#> other attached packages:
#> [1] LIPIDIFy_0.99.7  BiocStyle_2.41.0
#> 
#> loaded via a namespace (and not attached):
#>   [1] gridExtra_2.3.1             sandwich_3.1-3             
#>   [3] rlang_1.3.0                 magrittr_2.0.5             
#>   [5] shinydashboard_0.7.3        multcomp_1.4-32            
#>   [7] otel_0.2.0                  matrixStats_1.5.0          
#>   [9] compiler_4.6.1              vctrs_0.7.3                
#>  [11] stringr_1.6.0               sysfonts_0.8.9             
#>  [13] pkgconfig_2.0.3             fastmap_1.2.0              
#>  [15] XVector_0.53.0              magick_2.9.1               
#>  [17] labeling_0.4.3              promises_1.5.0             
#>  [19] rmarkdown_2.32              preprocessCore_1.75.1      
#>  [21] tinytex_0.61                purrr_1.2.2                
#>  [23] xfun_0.61                   showtext_0.9-8             
#>  [25] cachem_1.1.0                jsonlite_2.0.0             
#>  [27] flashClust_1.1-4            later_1.4.8                
#>  [29] DelayedArray_0.39.7         BiocParallel_1.47.0        
#>  [31] irlba_2.3.7                 parallel_4.6.1             
#>  [33] cluster_2.1.8.3             R6_2.6.1                   
#>  [35] bslib_0.12.0                stringi_1.8.9              
#>  [37] RColorBrewer_1.1-3          limma_3.99.0               
#>  [39] GenomicRanges_1.65.4        jquerylib_0.1.4            
#>  [41] estimability_2.0.0          Seqinfo_1.3.2              
#>  [43] SummarizedExperiment_1.43.0 Rcpp_1.1.2                 
#>  [45] bookdown_0.48               knitr_1.52                 
#>  [47] zoo_1.9-0                   IRanges_2.47.5             
#>  [49] httpuv_1.6.17               Matrix_1.7-6               
#>  [51] splines_4.6.1               tidyselect_1.2.1           
#>  [53] abind_1.4-8                 dichromat_2.0-1            
#>  [55] yaml_2.3.12                 affy_1.91.0                
#>  [57] codetools_0.2-20            lattice_0.23-1             
#>  [59] tibble_3.3.1                Biobase_2.73.2             
#>  [61] withr_3.0.3                 shiny_1.14.0               
#>  [63] S7_0.2.2                    coda_0.19-4.1              
#>  [65] evaluate_1.0.5              survival_3.8-12            
#>  [67] zip_3.0.2                   affyio_1.83.0              
#>  [69] pillar_1.11.1               BiocManager_1.30.27        
#>  [71] MatrixGenerics_1.25.0       stats4_4.6.1               
#>  [73] DT_0.34.0                   plotly_4.12.1              
#>  [75] generics_0.1.4              S4Vectors_0.51.10          
#>  [77] ggplot2_4.0.3               scales_1.4.0               
#>  [79] xtable_1.8-8                leaps_3.2                  
#>  [81] glue_1.8.1                  pheatmap_1.0.13            
#>  [83] emmeans_2.0.4               scatterplot3d_0.3-45       
#>  [85] tools_4.6.1                 data.table_1.18.6.1        
#>  [87] vsn_3.81.1                  openxlsx_4.2.9             
#>  [89] fgsea_1.39.2                locfit_1.5-9.12            
#>  [91] mvtnorm_1.4-2               fastmatch_1.1-8            
#>  [93] cowplot_1.2.0               grid_4.6.1                 
#>  [95] tidyr_1.3.2                 edgeR_4.99.6               
#>  [97] showtextdb_3.0              cli_3.6.6                  
#>  [99] S4Arrays_1.13.1             viridisLite_0.4.3          
#> [101] dplyr_1.2.1                 gtable_0.3.6               
#> [103] pls_2.9-0                   sass_0.4.10                
#> [105] digest_0.6.39               BiocGenerics_0.59.12       
#> [107] SparseArray_1.13.3          ggrepel_0.9.8              
#> [109] TH.data_1.1-5               FactoMineR_2.17            
#> [111] htmlwidgets_1.6.4           farver_2.1.2               
#> [113] htmltools_0.5.9             lifecycle_1.0.5            
#> [115] httr_1.4.9                  multcompView_0.1-12        
#> [117] statmod_1.5.2               mime_0.13                  
#> [119] MASS_7.3-66