Tutorial
The science behind oRNAment
Protein-RNA interactions execute important roles in several biological functions, including RNA replication, repair, splicing, polyadenylation, capping, modification, export, localization, stability and degradation. Recent in vitro technologies, such as RNAcompete [1] and RNA Bind-n-Seq (RBNS) [2], have provided exquisite resources for identifying the binding preferences of RBPs.
We acquired all motifs from RNAcompete in the form a position weight matrix (PWM) and executed the RBNS algorithm on the sequencing data made available by ENCODE. As most motifs determined by RNAcompete were of length 7, we concentrated on those and set the output motif PWM length of RBNS to also be 7. Therefore, all motifs in the database are of length 7 nucleotides and are comparable (see motifs tab).
We developed a novel algorithm that allows us to scan for these motifs with yet unachieved efficiency (see algorithm tab). We then executed it for each RBP on the complete coding and non-coding transcriptomes of human and 4 main model organisms described by Ensembl (C. elegans, D. rerio, D. melanogaster, M. musculus) in oRNAment v1, and further extended to 10 more model organisms in v2, bringing the total organisms scanned to 15. As there might by interesting evolutionary question answered with the help of this database, we scanned for each RBPs in all transcripts of every organism independently of whether a given RBP is expressed or not. We have shown that our method can accurately predict the putative binding sites observed by eCLIP [3] in humans (See validation tab).
[1] Ray, D., et al. (2013) A compendium of RNA-binding motifs for decoding gene regulation. Nature.
[2] Lambert,N., et al. (2014) RNA Bind-n-Seq: Quantitative Assessment of the Sequence and Structural Binding Specificity of RNA Binding Proteins. Mol Cell, 54, 887.
[3] Nostrand,E., et al. (2016) Robust transcriptome-wide discovery of RNA-binding protein binding sites with enhanced CLIP (eCLIP). Nat Methods, 13.
The search algorithm scans each transcript and for each subsequence of length 7, the length of each motif, calculate its matrix similarity score (MSS). The MSS is defined as MSS = current_score – mimimum_score / maximum_score – minimum score. This will provide a value between 0 and 1 where 1 is the canonical motif. The score_current is defined as the product of each probability to observe a given nucleotide at a given position in the PWM. The score_maximum is the product of each maximum probability value in the PWM at each position. The score_minimum is the product of each minimum probability value in the PWM at each position.
Independently, the sum of all MSS score for each 4^7 possible substrings (k-mer) was calculated. A score, MSS’, is calculated by sequentially adding the MSS value of each k-mer incrementally. A score MSS’ % is obtained by dividing, for each k-mer, its MSS’ by the sum of all MSS’. A threshold can thus be calculated by taking the MSS scores representing up to 50% of all possible scores, by steps of 5%.
As shown in the figure, when searching for a given RBP motif (A) in a specific transcript (B), we can rapidly scan for each k-mer (C) and directly look it up in a table (D) in constant time, through the use of a dictionary, and populate the database with substrings passing the threshold as putative motif instances (E).
In oRNAment V2 we decided to treat homologs from different species as different RBPs due to the increased number of species and predicted RBPs. In addition, we clustered these RBPs based on the similarity of their motifs. In general, all RBPs that share at least one motif across all species are clustered together. We further inspect the clustering to ensure that it is meaningful and name each cluster according to its most representative RBPs.
List of motifs and their logo. The logos and their Position Weight Matrix can be downloaded in the download page.
A1CF (1)
A1CF (10)
A1CF (2)
A1CF (3)
A1CF (4)
A1CF (5)
A1CF (6)
A1CF (7)
A1CF (8)
A1CF (9)
RBM46
RBM47 (1)
RBM47 (2)
ANKHD1
ARGLU1 (1)
ARGLU1 (2)
asd-1
B0280.17
BOLL (1)
BOLL (2)
BOLL (3)
BOLL (4)
C44B7.2
CELF (1)
CELF (2)
CELF (3)
CELF (4)
CELF (5)
CELF (6)
CELF (7)
CELF1 (1)
CELF1 (2)
CELF1 (3)
CG1316
CG15440
CG33714
CG4887
CG7804
CG7903
CIRBP
CNOT4 (1)
CNOT4 (2)
CNOT4 (3)
CNOT4 (4)
CNOT4 (5)
CNOT4 (6)
cpb-2
cpb-3
CPEB2
CPEB3 (1)
CPEB3 (2)
CPEB1 (2)
CPEB1 (1)
DAZ3 (1)
DAZ3 (2)
DAZ3 (3)
DAZ3 (4)
DAZ3 (5)
DAZ3 (6)
DAZAP1 (1)
DAZAP1 (2)
DAZAP1 (3)
DAZAP1 (4)
DDX53
EIF2S1
EIF4G2 (1)
EIF4G2 (2)
EIF4G2 (3)
EIF4G2 (4)
ELAVL1 (1)
ELAVL1 (2)
ELAVL1 (3)
ELAVL1 (4)
ELAVL1 (5)
ELAVL2 (1)
ELAVL2 (2)
ELAVL4 (1)
ELAVL4 (2)
ELAVL4 (3)
elav
ENOX1
ENSXETG00000023721
ESRP1 (1)
ESRP1 (2)
ESRP1 (3)
ESRP2
etr-1
EWSR1 (1)
EWSR1 (2)
EWSR1 (3)
EWSR1 (4)
exc-7
Fmr1
FMR1 (1)
FMR1 (2)
FXR2
A4QNI8
FUBP1 (1)
FUBP1 (2)
FUBP1 (3)
FUBP1 (4)
FUBP3 (1)
FUBP3 (2)
FUS (1)
FUS (2)
G3BP2
H28G03.1
heph
HNRNPA1
HNRNPA1L2
HNRNPA2B1 (1)
HNRNPA2B1 (2)
HNRNPA2B1 (3)
HNRNPA0 (1)
HNRNPA0 (2)
HNRNPA0 (3)
HNRNPA0 (4)
HNRNPA0 (5)
HNRNPA0 (6)
HNRNPA0 (7)
HNRNPA0 (8)
HNRNPAB
HNRNPD (1)
HNRNPD (2)
HNRNPD (3)
HNRNPD (4)
HNRNPDL (1)
HNRNPDL (2)
HNRNPF (1)
HNRNPF (2)
HNRNPF (3)
HNRNPH1
HNRNPH2 (1)
HNRNPH2 (2)
HNRNPK (1)
HNRNPK (2)
HNRNPK (3)
HNRNPK (4)
HNRNPK (5)
HNRNPL (1)
HNRNPL (2)
HNRNPL (3)
HNRNPL (4)
HNRNPLL
Hrb27C (1)
Hrb27C (2)
Hrb87F
Hrb98DE (1)
Hrb98DE (2)
Hrb98DE (3)
HRP1
hrp-2
IGF2BP1 (1)
IGF2BP1 (2)
IGF2BP1 (3)
IGF2BP2 (1)
IGF2BP2 (2)
IGF2BP2 (3)
IGF2BP2 (4)
IGF2BP2 (5)
IGF2BP3
ILF2 (1)
ILF2 (2)
ILF2 (3)
Imp
K07H8.9
KHDRBS1 (1)
KHDRBS1 (2)
KHDRBS1 (3)
KHDRBS2 (1)
KHDRBS2 (2)
KHDRBS2 (3)
KHDRBS2 (4)
KHDRBS3 (1)
KHDRBS3 (2)
KHDRBS3 (3)
KHDRBS3 (4)
KHSRP (1)
KHSRP (2)
lark (1)
lark (2)
LIN28A (1)
LIN28A (2)
lin-41
MATR3
MBNL1 (1)
MBNL1 (2)
MBNL1 (3)
MBNL1 (4)
MBNL1 (5)
MBNL2
mec-8
mei-P26
mex3c
mod
MSI1 (1)
MSI1 (2)
MSI1 (3)
MSI1 (4)
MSI1 (5)
msi (1)
msi (2)
msi (3)
mub
mxt
Nab2
ncl-1
ENSG00000206268
nhl-2
NOVA1 (1)
NOVA1 (2)
NOVA1 (3)
NOVA1 (4)
NUDT21
NUPL2 (1)
NUPL2 (2)
NUPL2 (3)
NUPL2 (4)
NUPL2 (5)
NUPL2 (6)
NUPL2 (7)
pab-1
PABPC1
PABPC1L (1)
PABPC1L (2)
PABPC1L2A
PABPC4
PABPC5
BCL2L2-PABPN1
PABPN1L (1)
PABPN1L (2)
papi
PCBP1 (1)
PCBP1 (2)
PCBP1 (3)
PCBP1 (4)
PCBP1 (5)
PCBP2 (1)
PCBP2 (2)
PCBP2 (3)
PCBP2 (4)
PCBP2 (5)
PCBP2 (6)
PCBP3
PCBP4 (1)
PCBP4 (2)
PCBP4 (3)
PPRC1
PRR3 (4)
PRR3 (1)
PRR3 (2)
PRR3 (3)
PSPC1
SFPQ (1)
SFPQ (2)
SFPQ (3)
SFPQ (4)
SFPQ (5)
SFPQ (6)
PTBP1 (1)
PTBP1 (2)
PTBP2
PTBP3 (1)
PTBP3 (2)
PUB1
PUF60 (3)
PUF60 (4)
PUF60 (1)
PUF60 (2)
pUf68
PUM1 (1)
PUM1 (2)
PUM1 (3)
PUM1 (4)
PUM1 (5)
PUM1 (6)
PUM1 (7)
PUM1 (8)
QKI (1)
QKI (2)
CG3927
RALY;HNRNPC (1)
RALY;HNRNPC (2)
RALY;HNRNPC (3)
HNRNPC (1)
HNRNPC (2)
HNRNPCL1 (1)
HNRNPCL1 (2)
RALYL (1)
RALYL (2)
Rb97D
RBFOX (1)
RBFOX (2)
RBFOX (3)
RBFOX2
RBFOX3
RBM10 (1)
RBM10 (2)
RBM15B (4)
RBM15B (1)
RBM15B (2)
RBM15B (3)
RBM22
RBM23 (1)
RBM23 (2)
RBM24/38 (1)
RBM24/38 (2)
RBM24/38 (3)
RBM24 (1)
RBM24 (2)
RBM24 (3)
RBM38
sup-12
RBM25 (1)
RBM25 (2)
RBM25 (3)
RBM28
RBM34
RBM14 (1)
RBM14 (2)
RBM4 (1)
RBM4 (2)
RBM4 (3)
RBM4 (4)
RBM4 (5)
RBM4B (1)
RBM4B (2)
RBM4B (3)
RBM4B (4)
RBM41 (1)
RBM41 (2)
RBM41 (3)
RBM42 (1)
RBM42 (2)
RBM45 (1)
RBM45 (2)
RBM45 (3)
RBM45 (4)
RBM45 (5)
RBM45 (6)
RBM45 (7)
RBM45 (8)
RBM5 (1)
RBM5 (2)
RBM6 (1)
RBM6 (2)
RBM6 (3)
RBM6 (4)
RBM6 (5)
RBM8A
RBMS1
RBMS2 (1)
RBMS2 (2)
RBMS2 (3)
RBMS3 (1)
RBMS3 (2)
RBMS3 (3)
RBMS3 (4)
RBMS3 (5)
RBMS3 (6)
RBPMS (1)
RBPMS (2)
RC3H1 (1)
RC3H1 (2)
RC3H1 (3)
RC3H1 (4)
Ref2
rnp-2
Rnp4F
RNPC3
Rox8
Rsf1
rsp-1
rsp-2
SAMD4A
SART3
SERBP1
SF1 (1)
SF1 (2)
SF1 (3)
SF1 (4)
SF3B4 (1)
SF3B4 (2)
SF3B6
sfa-1
shep (1)
shep (2)
shep (3)
sm
SNRNP70
SNRPA (1)
SNRPA (2)
SNRPA (3)
SNRPA (4)
sqd-1
SRSF1 (1)
SRSF1 (2)
SRSF1 (3)
SRSF1 (4)
SRSF1 (5)
SRSF1 (6)
SRSF9 (1)
SRSF9 (2)
SRSF9 (3)
SRSF9 (4)
SRSF9 (5)
srsf1a
SRSF10 (1)
SRSF10 (2)
SRSF10 (3)
SRSF10 (4)
SRSF10 (5)
SRSF10 (6)
SRSF11 (1)
SRSF11 (2)
SRSF2 (1)
SRSF2 (2)
SRSF2 (3)
SRSF2 (4)
Rbp1
Rbp1-like
SRSF3 (1)
SRSF3 (2)
SRSF4 (1)
SRSF4 (2)
SRSF4 (3)
SRSF6
SRSF5 (1)
SRSF5 (2)
SRSF5 (3)
SRSF5 (4)
SRSF5 (5)
SRSF8 (1)
SRSF8 (2)
ssx
sup-26
HNRNPR (1)
HNRNPR (2)
HNRNPR (3)
Syp
T07F10.3
TAF15 (1)
TAF15 (2)
TAF15 (3)
TARDBP (3)
TARDBP (4)
TARDBP (1)
TARDBP (2)
tdp-1
TIA1 (1)
TIA1 (2)
TIA1 (3)
TIA1 (4)
TIA1 (5)
TIA1 (6)
tiar-1
tiar-3
tra2
TRA2A (1)
TRA2A (2)
TRIM56
TRIM71
TRNAU1AP (1)
TRNAU1AP (2)
TRNAU1AP (3)
TRNAU1AP (4)
TUT1
U2AF2 (1)
U2AF2 (2)
U2AF2 (3)
U2AF2 (4)
UNK (1)
UNK (2)
UNK (3)
UNK (4)
wech
Y57G11C.36
Y59A8B.10
YBX1 (1)
YBX1 (2)
YBX2
ZC3H10
ZC3H14
CNBP
ZCRB1 (1)
ZCRB1 (2)
ZFP36
ZNF326 (1)
ZNF326 (2)
ZNF326 (3)
ZNF326 (4)
ZNF638
To determine the age of motif hits across the 7 mammalian species, we identified the corresponding ancestral nodes for each of our species from the phylogeny derived from the Zoonomia consortium alignment of 447 mammalian genomes (Kuderna et al., Nature, 2023; Genereux et al., Nature, 2020; Cock et al., Biopython, 2009). We estimated the approximate age of the ancestral sequences from the TimeTree database (Kumar et al., MBE, 2022). We used halliftover to map motif-hit coordinates from the 7 mammals to their respective inferred ancestral sequences (Hickey et al., Bioinformatics, 2013). With bedtools, we then extended the lifted over coordinates by 6 bp on either side and extracted the corresponding ancestral sequences (Quinlan and Hall, Bioinformatics, 2010). We scanned the resulting sequences for their respective pre-calculated motif, and classified a sequence as a hit if its MSS score was at least in the 50th percentile. Finally, we estimated the approximate age of a motif hit by the age of the oldest ancestor for which the mapped region had a motif hit, with unmapped sequences classified as of Human origin.
The genomic region of motif instances identified by oRNAment at a threshold of 50% greatly correspond to the binding regions observed by eCLIP at a threshold of 3 fold change and a p-value of 0.001. Furthermore, the same number of random genomic regions as observed by oRNAment are seldom matching a binding region in eCLIP.
In oRNAment V2 we improved our database in the following aspects:
- Motif instances are now predicted using the recent Ensembl release 115 (September 2025).
-
Ten new species have been added, bringing the total to 15 species covered in the database:
- Human (Homo sapiens) Genome Assembly GRCh38
- Mouse (Mus musculus) Genome Assembly GRCm39
- Fruit Fly (Drosophila melanogaster) Genome Assembly BDGP6.54
- Zebrafish (Danio rerio) Genome Assembly GRCz11
- Roundworm (Caenorhabditis elegans) Genome Assembly WBcel235
- Chimpanzee (Pan troglodytes) Genome Assembly Pan_tro_3.0
- Rhesus macaque (Macaca mulatta) Genome Assembly Mmul_10
- Olive baboon (Papio anubis) Genome Assembly Panubis1.0
- Gray mouse lemur (Microcebus murinus) Genome Assembly Mmur_3.0
- Rat (Rattus norvegicus) Genome Assembly GRCr8
- Chicken (Gallus gallus) Genome Assembly bGalGal1.mat.broiler.GRCg7b
- Western clawed frog (Xenopus tropicalis) Genome Assembly UCB_Xtro_10.0
- Budding yeast (Saccharomyces cerevisiae) Genome Assembly R64-1-1
- Lizard (Anolis carolinensis) Genome Assembly AnoCar2.0v2
- Zebra finch (Taeniopygia guttata) Genome Assembly bTaeGut1_v1.p
- Expanded RBP and motif coverage: the database now includes 2,089 RBPs across 15 species and 492 motifs. Homologous RBPs from different species are counted separately, though they typically share the same motif. The complete relationship table between the motifs and the RBPs can be found in the download page
- Extended scan regions to include up to 150 nt of intronic sequence flanking exons, enabling coverage of splicing regulatory signals and other functions dependent on exon–intron boundaries. Intron are tagged through their host transcript, in the format of (transcript id)_(intron number)_(side), where side indicates whether it involves the 5' or 3' of the intron, or 5-3 if the entire intron is shorter than 300 nt.
- Added search function to table views of motif instances to help navigate through all the motif instances identified in our database.
- Added links to UCSC genome browser track hubs to interactively browse motif instances
- Added estimated motif age for mammals to show the conservation of motifs
Please cite:
Louis Philip Benoit Bouvrette, Samantha Bovaird, Mathieu Blanchette, Eric Lécuyer oRNAment: A database of putative RNA binding protein target sites in the transcriptomes of model species. https://doi.org/10.1093/nar/gkz986This site has been validated on Chrome, Firefox, Safari, and Vivaldi on both macOS and Windows environments. For best user experience we recommend Chrome or Vivaldi. Some functionality may not work if using Explorer on Windows.
The search by... forms
This tool allows you to search the database for a specific gene or transcript, for which you would like to know all putative RBP motif instances (by leaving the RBP choice as "All RBPs"), or for one RBP by choosing it from the drop down list. The choice of RBP also include the species to which it belongs, so the species-specific flag is ignored.
All choice fields are searchable.
In v2 only one gene is allowed at a time to reduce server pressure as by including intron and expanding our RBP list, a lot more motifs are detected than in the previous version.
This tool allows you to search the database for a specific combination of attributes in a given species for which you would like to know all putative RBP motif instances.
Select the species and the biotype or region of interest. Notice that not all combinations of choices are compatible.
This tool allows you to search the database for a specific RBP from a specific species, at a specific similarity threshold between the PWM and the subsequence, for which you would like to know all its putative instances in all the coding and non-coding transcripts of the selected organism.
Select your organism and your RBP of interest. Input the threshold for motif score, whose range is limited between 0.5 and 1 to limit output table size while allowing for good overlap between motif prediction and in vitro binding essay results.
All choice fields are searchable.
The results pages
The different search tools provide for various top-level graph analysis and a detailed table of all motif instances in each transcript. Kindly note that, by design, the database uses a "transcript-centric" view. Motifs instances coming from different isoform of a gene may have the same genomic coordinates. Therefore, a motif instance can be present more than once in the table, but will always represent two or more distinct transcripts of a gene.
The result page starts always with a call back of search criteria
The following plots and tables are interactive and searchable when a lot of entries are expected.
The data for each plot and entire table of motif instances can be downloaded as file. Please note that the CSV may take a moment to generate.
When searching by gene name it could map to multiple Ensembl Gene ID if they are found on alternative contigs among other reasons. When searching by transcript, while we executed the scan on all transcripts available from the Ensembl database, due to varying nomenclature and versioning, it is possible that a specific gene name/ID will not be found in oRNAment.
Either case a table containing identified gene and transcripts is shown in the result page.
The first plot shows the motif abundance through bar plot, histogram, tree map and searchable table. Different motifs of the same RBP are displayed individually.
The next plot shows the motif distribution across transcript regions. Please note that the entire squence of non-coding trancripts from a coding gene is classified as "Non Coding" despite it may share identical regions to UTRs from its coding counter parts. However since the "coding region" is technically undefined for these transcripts, it is not possible to distinguish UTR from the rest of the transcript.
Finally, all motif instances are shown as a table where it is possible to filter each column to facilitate browsing.
Please notice that not all combinations of search criteria will produce results. For example, choosing "snoRNA" for gene and transcript types while "CDS" for transcript region is not a valid combination, as snoRNAs do not have CDS. We are actively working to handle the selection logic to eleminate non-valid combinations.
The first plot shows the motif abundance through bar plot, histogram, tree map and searchable table. Different motifs of the same RBP are displayed individually.
Due to potential high row count of the motif instances, we do not offer browserable view of the table. A downloadable csv is still offered, with the total number of rows to help verifying the completeness of file.
The first plot shows the biotypes of gene harboring the motif(s) of the RBP through bar plot, histogram, tree map and searchable table.
The next plot shows the motif distribution across transcript regions. Please note that the entire squence of non-coding trancripts from a coding gene is classified as "Non Coding" despite it may share identical regions to UTRs from its coding counter parts. However since the "coding region" is technically undefined for these transcripts, it is not possible to distinguish UTR from the rest of the transcript.
Due to potential high row count of the motif instances, we do not offer browserable view of the table. A downloadable csv is still offered, with the total number of rows to help verifying the completion of file.
After the call back, all motifs from the same cluster (see above) are displayed with a marker shape assigned to each of them.
This is followed by the plots for transcripts and introns from the given gene or transcript. If a transcript or an intron does not contain any motifs from the motif cluster, the transcript or intron is omitted from this page.
On the left, the motifs are plotted with the transcript position as x-axis, and the motif score as y-axis, with the shade indicating its unpaired probability. On the right, we show the structure predicted by RNAfold and visualized by FornaContainer.
We decided to not predict structure for the intronic sequences, as it is unlikely to reflect the true structure of the intron in its context. In this case the unpaired probability is set to 1 solely for the purpose of display for introns.
The feature ID of an intron is derived from its host transcript, in the format of (transcript id)_(intron number)_(side), where side indicates whether it involves the 5' or 3' of the intron, or 5-3 if the entire intron is shorter than 300 nt. We do not include predicted structure for introns as the predictions based on 150 nt region are unlikely to reflect the true structure of introns within their context.