| Nailpolish version | nailpolish v0.2.2, commit #65a177c-modified |
| File path | scmixology2_sample.fastq |
| Dataset size | 0.02817634679377079 GB |
| Index date | 2026-08-13T16:17:08.927358+10:00 |
| Total read count | 14143 |
| Reads with barcodes | 14143 |
| Reads without barcodes (ignored from consensus calling) |
0 |
| Average quality | 21.152048110961914 |
| Average length | 1030.5797119140625 |
A 'duplicate group' is a group of reads which all share the same barcode and UMI.
There are calculating... total duplicate group(s). Only groups with < duplicates are shown in the graph.
Note. Some reads will share a molecular identifier despite originating from different molecules (i.e. false duplicates). They are represented together as duplicate groups in this chart. Nailpolish's false clustering algorithm will be applied during consensus calling to identify and separate these reads into distinct groups. As a result, the final number of groups (and their size distribution) in the consensus calling output may differ.
Each read is classified by the number of reads in its corresponding duplicate group.
There are calculating... total reads. Only reads belonging to groups with < duplicates are shown in the graph.
| Group Size | Number of Groups | Number of Reads | Average Length | Average Quality |
|---|
This table shows the full distribution of duplicate group sizes across the dataset.
Each row corresponds to a group size (number of reads sharing the same barcode and UMI), along with the number of such groups, total reads, and average read length and quality within that group size.
--len (default 0,15000): does not consensus call reads with length outside the
specified range; these reads are reproduced as-is in the consensus calling output
--qual (default unset): does not consensus call reads with mean quality outside the
specified range; these reads are reproduced as-is in the consensus calling output
--max-group-size (default 250): does not consensus call reads belonging to duplicate
groups larger than the specified size. Typically, very large groups are false duplicates and
consensus calling them has an outsized impact on runtime. These reads are reproduced as-is in the
consensus
calling
output