Intervals and layouts
Each interval combines a role prefix with a shape.
Interval roles
Section titled “Interval roles”| Prefix | Role | Typical use |
|---|---|---|
b |
barcode | Cell, round, or other identifying barcode. |
s |
sample barcode | Barcode used to distinguish samples. |
u |
UMI | Unique molecular identifier. |
r |
read sequence | Biological sequence retained for downstream analysis. |
x |
discard | Sequence that is consumed but not needed in output. |
f |
fixed sequence | Literal adapter, linker, or anchor sequence. |
The role records intent and lets the compiler build named sequence components; it does not by itself correct or filter an interval.
Interval shapes
Section titled “Interval shapes”| Shape | Meaning | Example |
|---|---|---|
[n] |
Exactly n bases. |
u[10] |
[a-b] |
Between a and b bases, inclusive. |
b[9-10] |
: |
The remaining unbounded sequence. | r: |
[ACGT] |
A literal nucleotide sequence; valid for f. |
f[CAGAGC] |
Variable and unbounded intervals need a surrounding layout that makes their boundaries determinable. An anchor is a common way to delimit a variable prefix.
Read layouts
Section titled “Read layouts”Braces group intervals belonging to an input read:
1{b[16]u[10]}2{r:}Paired files are supplied as --file1 and --file2. The output transformation
uses the same numbered layout notation:
1{b<bc>[16]u<umi>[10]}2{r<bio>:}-> 1{<bc><umi>} 2{<bio>}An output can contain only extracted labels; it need not reproduce consumed anchors or discarded sequence.
Definitions and references
Section titled “Definitions and references”Prefer definitions when an interval has a protocol-level name or operation:
bc = filter_within_dist(b[8], "barcodes.txt", 1)umi = u[10]1{<bc><umi>r:}-> 1{<bc><umi>}Definitions also provide the attachment point for operation annotations such
as #[edit(1)] and the property annotation #[ambig_policy = no_match].