Tutorial

The science behind oRNAment

Protein-RNA interactions execute important roles in several biological functions, including RNA replication, repair, splicing, polyadenylation, capping, modification, export, localization, stability and degradation. Recent in vitro technologies, such as RNAcompete [1] and RNA Bind-n-Seq (RBNS) [2], have provided exquisite resources for identifying the binding preferences of RBPs.

We acquired all motifs from RNAcompete in the form a position weight matrix (PWM) and executed the RBNS algorithm on the sequencing data made available by ENCODE. As most motifs determined by RNAcompete were of length 7, we concentrated on those and set the output motif PWM length of RBNS to also be 7. Therefore, all motifs in the database are of length 7 nucleotides and are comparable (see motifs tab).

We developed a novel algorithm that allows us to scan for these motifs with yet unachieved efficiency (see algorithm tab). We then executed it for each RBP on the complete coding and non-coding transcriptomes of human and 4 main model organisms described by Ensembl (C. elegans, D. rerio, D. melanogaster, M. musculus) in oRNAment v1, and further extended to 10 more model organisms in v2, bringing the total organisms scanned to 15. As there might by interesting evolutionary question answered with the help of this database, we scanned for each RBPs in all transcripts of every organism independently of whether a given RBP is expressed or not. We have shown that our method can accurately predict the putative binding sites observed by eCLIP [3] in humans (See validation tab).

[1] Ray, D., et al. (2013) A compendium of RNA-binding motifs for decoding gene regulation. Nature.
[2] Lambert,N., et al. (2014) RNA Bind-n-Seq: Quantitative Assessment of the Sequence and Structural Binding Specificity of RNA Binding Proteins. Mol Cell, 54, 887.
[3] Nostrand,E., et al. (2016) Robust transcriptome-wide discovery of RNA-binding protein binding sites with enhanced CLIP (eCLIP). Nat Methods, 13.

The search algorithm scans each transcript and for each subsequence of length 7, the length of each motif, calculate its matrix similarity score (MSS). The MSS is defined as MSS = current_score – mimimum_score / maximum_score – minimum score. This will provide a value between 0 and 1 where 1 is the canonical motif. The score_current is defined as the product of each probability to observe a given nucleotide at a given position in the PWM. The score_maximum is the product of each maximum probability value in the PWM at each position. The score_minimum is the product of each minimum probability value in the PWM at each position.

Independently, the sum of all MSS score for each 4^7 possible substrings (k-mer) was calculated. A score, MSS’, is calculated by sequentially adding the MSS value of each k-mer incrementally. A score MSS’ % is obtained by dividing, for each k-mer, its MSS’ by the sum of all MSS’. A threshold can thus be calculated by taking the MSS scores representing up to 50% of all possible scores, by steps of 5%.

As shown in the figure, when searching for a given RBP motif (A) in a specific transcript (B), we can rapidly scan for each k-mer (C) and directly look it up in a table (D) in constant time, through the use of a dictionary, and populate the database with substrings passing the threshold as putative motif instances (E).

Science_Math_Algo.png

In oRNAment V2 we decided to treat homologs from different species as different RBPs due to the increased number of species and predicted RBPs. In addition, we clustered these RBPs based on the similarity of their motifs. In general, all RBPs that share at least one motif across all species are clustered together. We further inspect the clustering to ensure that it is meaningful and name each cluster according to its most representative RBPs.

List of motifs and their logo. The logos and their Position Weight Matrix can be downloaded in the download page.

A1CF (1)

A1CF (1)

A1CF (10)

A1CF (10)

A1CF (2)

A1CF (2)

A1CF (3)

A1CF (3)

A1CF (4)

A1CF (4)

A1CF (5)

A1CF (5)

A1CF (6)

A1CF (6)

A1CF (7)

A1CF (7)

A1CF (8)

A1CF (8)

A1CF (9)

A1CF (9)

RBM46

RBM46

RBM47 (1)

RBM47 (1)

RBM47 (2)

RBM47 (2)

ANKHD1

ANKHD1

ARGLU1 (1)

ARGLU1 (1)

ARGLU1 (2)

ARGLU1 (2)

asd-1

asd-1

B0280.17

B0280.17

BOLL (1)

BOLL (1)

BOLL (2)

BOLL (2)

BOLL (3)

BOLL (3)

BOLL (4)

BOLL (4)

C44B7.2

C44B7.2

CELF  (1)

CELF (1)

CELF  (2)

CELF (2)

CELF  (3)

CELF (3)

CELF  (4)

CELF (4)

CELF  (5)

CELF (5)

CELF  (6)

CELF (6)

CELF  (7)

CELF (7)

CELF1 (1)

CELF1 (1)

CELF1 (2)

CELF1 (2)

CELF1 (3)

CELF1 (3)

CG1316

CG1316

CG15440

CG15440

CG33714

CG33714

CG4887

CG4887

CG7804

CG7804

CG7903

CG7903

CIRBP

CIRBP

CNOT4 (1)

CNOT4 (1)

CNOT4 (2)

CNOT4 (2)

CNOT4 (3)

CNOT4 (3)

CNOT4 (4)

CNOT4 (4)

CNOT4 (5)

CNOT4 (5)

CNOT4 (6)

CNOT4 (6)

cpb-2

cpb-2

cpb-3

cpb-3

CPEB2

CPEB2

CPEB3 (1)

CPEB3 (1)

CPEB3 (2)

CPEB3 (2)

CPEB1 (2)

CPEB1 (2)

CPEB1 (1)

CPEB1 (1)

DAZ3 (1)

DAZ3 (1)

DAZ3 (2)

DAZ3 (2)

DAZ3 (3)

DAZ3 (3)

DAZ3 (4)

DAZ3 (4)

DAZ3 (5)

DAZ3 (5)

DAZ3 (6)

DAZ3 (6)

DAZAP1 (1)

DAZAP1 (1)

DAZAP1 (2)

DAZAP1 (2)

DAZAP1 (3)

DAZAP1 (3)

DAZAP1 (4)

DAZAP1 (4)

DDX53

DDX53

EIF2S1

EIF2S1

EIF4G2 (1)

EIF4G2 (1)

EIF4G2 (2)

EIF4G2 (2)

EIF4G2 (3)

EIF4G2 (3)

EIF4G2 (4)

EIF4G2 (4)

ELAVL1 (1)

ELAVL1 (1)

ELAVL1 (2)

ELAVL1 (2)

ELAVL1 (3)

ELAVL1 (3)

ELAVL1 (4)

ELAVL1 (4)

ELAVL1 (5)

ELAVL1 (5)

ELAVL2 (1)

ELAVL2 (1)

ELAVL2 (2)

ELAVL2 (2)

ELAVL4 (1)

ELAVL4 (1)

ELAVL4 (2)

ELAVL4 (2)

ELAVL4 (3)

ELAVL4 (3)

elav

elav

ENOX1

ENOX1

ENSXETG00000023721

ENSXETG00000023721

ESRP1 (1)

ESRP1 (1)

ESRP1 (2)

ESRP1 (2)

ESRP1 (3)

ESRP1 (3)

ESRP2

ESRP2

etr-1

etr-1

EWSR1 (1)

EWSR1 (1)

EWSR1 (2)

EWSR1 (2)

EWSR1 (3)

EWSR1 (3)

EWSR1 (4)

EWSR1 (4)

exc-7

exc-7

Fmr1

Fmr1

FMR1 (1)

FMR1 (1)

FMR1 (2)

FMR1 (2)

FXR2

FXR2

A4QNI8

A4QNI8

FUBP1 (1)

FUBP1 (1)

FUBP1 (2)

FUBP1 (2)

FUBP1 (3)

FUBP1 (3)

FUBP1 (4)

FUBP1 (4)

FUBP3 (1)

FUBP3 (1)

FUBP3 (2)

FUBP3 (2)

FUS (1)

FUS (1)

FUS (2)

FUS (2)

G3BP2

G3BP2

H28G03.1

H28G03.1

heph

heph

HNRNPA1

HNRNPA1

HNRNPA1L2

HNRNPA1L2

HNRNPA2B1 (1)

HNRNPA2B1 (1)

HNRNPA2B1 (2)

HNRNPA2B1 (2)

HNRNPA2B1 (3)

HNRNPA2B1 (3)

HNRNPA0 (1)

HNRNPA0 (1)

HNRNPA0 (2)

HNRNPA0 (2)

HNRNPA0 (3)

HNRNPA0 (3)

HNRNPA0 (4)

HNRNPA0 (4)

HNRNPA0 (5)

HNRNPA0 (5)

HNRNPA0 (6)

HNRNPA0 (6)

HNRNPA0 (7)

HNRNPA0 (7)

HNRNPA0 (8)

HNRNPA0 (8)

HNRNPAB

HNRNPAB

HNRNPD (1)

HNRNPD (1)

HNRNPD (2)

HNRNPD (2)

HNRNPD (3)

HNRNPD (3)

HNRNPD (4)

HNRNPD (4)

HNRNPDL (1)

HNRNPDL (1)

HNRNPDL (2)

HNRNPDL (2)

HNRNPF (1)

HNRNPF (1)

HNRNPF (2)

HNRNPF (2)

HNRNPF (3)

HNRNPF (3)

HNRNPH1

HNRNPH1

HNRNPH2 (1)

HNRNPH2 (1)

HNRNPH2 (2)

HNRNPH2 (2)

HNRNPK (1)

HNRNPK (1)

HNRNPK (2)

HNRNPK (2)

HNRNPK (3)

HNRNPK (3)

HNRNPK (4)

HNRNPK (4)

HNRNPK (5)

HNRNPK (5)

HNRNPL (1)

HNRNPL (1)

HNRNPL (2)

HNRNPL (2)

HNRNPL (3)

HNRNPL (3)

HNRNPL (4)

HNRNPL (4)

HNRNPLL

HNRNPLL

Hrb27C (1)

Hrb27C (1)

Hrb27C (2)

Hrb27C (2)

Hrb87F

Hrb87F

Hrb98DE (1)

Hrb98DE (1)

Hrb98DE (2)

Hrb98DE (2)

Hrb98DE (3)

Hrb98DE (3)

HRP1

HRP1

hrp-2

hrp-2

IGF2BP1 (1)

IGF2BP1 (1)

IGF2BP1 (2)

IGF2BP1 (2)

IGF2BP1 (3)

IGF2BP1 (3)

IGF2BP2 (1)

IGF2BP2 (1)

IGF2BP2 (2)

IGF2BP2 (2)

IGF2BP2 (3)

IGF2BP2 (3)

IGF2BP2 (4)

IGF2BP2 (4)

IGF2BP2 (5)

IGF2BP2 (5)

IGF2BP3

IGF2BP3

ILF2 (1)

ILF2 (1)

ILF2 (2)

ILF2 (2)

ILF2 (3)

ILF2 (3)

Imp

Imp

K07H8.9

K07H8.9

KHDRBS1 (1)

KHDRBS1 (1)

KHDRBS1 (2)

KHDRBS1 (2)

KHDRBS1 (3)

KHDRBS1 (3)

KHDRBS2 (1)

KHDRBS2 (1)

KHDRBS2 (2)

KHDRBS2 (2)

KHDRBS2 (3)

KHDRBS2 (3)

KHDRBS2 (4)

KHDRBS2 (4)

KHDRBS3 (1)

KHDRBS3 (1)

KHDRBS3 (2)

KHDRBS3 (2)

KHDRBS3 (3)

KHDRBS3 (3)

KHDRBS3 (4)

KHDRBS3 (4)

KHSRP (1)

KHSRP (1)

KHSRP (2)

KHSRP (2)

lark (1)

lark (1)

lark (2)

lark (2)

LIN28A (1)

LIN28A (1)

LIN28A (2)

LIN28A (2)

lin-41

lin-41

MATR3

MATR3

MBNL1 (1)

MBNL1 (1)

MBNL1 (2)

MBNL1 (2)

MBNL1 (3)

MBNL1 (3)

MBNL1 (4)

MBNL1 (4)

MBNL1 (5)

MBNL1 (5)

MBNL2

MBNL2

mec-8

mec-8

mei-P26

mei-P26

mex3c

mex3c

mod

mod

MSI1 (1)

MSI1 (1)

MSI1 (2)

MSI1 (2)

MSI1 (3)

MSI1 (3)

MSI1 (4)

MSI1 (4)

MSI1 (5)

MSI1 (5)

msi (1)

msi (1)

msi (2)

msi (2)

msi (3)

msi (3)

mub

mub

mxt

mxt

Nab2

Nab2

ncl-1

ncl-1

ENSG00000206268

ENSG00000206268

nhl-2

nhl-2

NOVA1 (1)

NOVA1 (1)

NOVA1 (2)

NOVA1 (2)

NOVA1 (3)

NOVA1 (3)

NOVA1 (4)

NOVA1 (4)

NUDT21

NUDT21

NUPL2 (1)

NUPL2 (1)

NUPL2 (2)

NUPL2 (2)

NUPL2 (3)

NUPL2 (3)

NUPL2 (4)

NUPL2 (4)

NUPL2 (5)

NUPL2 (5)

NUPL2 (6)

NUPL2 (6)

NUPL2 (7)

NUPL2 (7)

pab-1

pab-1

PABPC1

PABPC1

PABPC1L (1)

PABPC1L (1)

PABPC1L (2)

PABPC1L (2)

PABPC1L2A

PABPC1L2A

PABPC4

PABPC4

PABPC5

PABPC5

BCL2L2-PABPN1

BCL2L2-PABPN1

PABPN1L (1)

PABPN1L (1)

PABPN1L (2)

PABPN1L (2)

papi

papi

PCBP1 (1)

PCBP1 (1)

PCBP1 (2)

PCBP1 (2)

PCBP1 (3)

PCBP1 (3)

PCBP1 (4)

PCBP1 (4)

PCBP1 (5)

PCBP1 (5)

PCBP2 (1)

PCBP2 (1)

PCBP2 (2)

PCBP2 (2)

PCBP2 (3)

PCBP2 (3)

PCBP2 (4)

PCBP2 (4)

PCBP2 (5)

PCBP2 (5)

PCBP2 (6)

PCBP2 (6)

PCBP3

PCBP3

PCBP4 (1)

PCBP4 (1)

PCBP4 (2)

PCBP4 (2)

PCBP4 (3)

PCBP4 (3)

PPRC1

PPRC1

PRR3 (4)

PRR3 (4)

PRR3 (1)

PRR3 (1)

PRR3 (2)

PRR3 (2)

PRR3 (3)

PRR3 (3)

PSPC1

PSPC1

SFPQ (1)

SFPQ (1)

SFPQ (2)

SFPQ (2)

SFPQ (3)

SFPQ (3)

SFPQ (4)

SFPQ (4)

SFPQ (5)

SFPQ (5)

SFPQ (6)

SFPQ (6)

PTBP1 (1)

PTBP1 (1)

PTBP1 (2)

PTBP1 (2)

PTBP2

PTBP2

PTBP3 (1)

PTBP3 (1)

PTBP3 (2)

PTBP3 (2)

PUB1

PUB1

PUF60 (3)

PUF60 (3)

PUF60 (4)

PUF60 (4)

PUF60 (1)

PUF60 (1)

PUF60 (2)

PUF60 (2)

pUf68

pUf68

PUM1 (1)

PUM1 (1)

PUM1 (2)

PUM1 (2)

PUM1 (3)

PUM1 (3)

PUM1 (4)

PUM1 (4)

PUM1 (5)

PUM1 (5)

PUM1 (6)

PUM1 (6)

PUM1 (7)

PUM1 (7)

PUM1 (8)

PUM1 (8)

QKI (1)

QKI (1)

QKI (2)

QKI (2)

CG3927

CG3927

RALY;HNRNPC (1)

RALY;HNRNPC (1)

RALY;HNRNPC (2)

RALY;HNRNPC (2)

RALY;HNRNPC (3)

RALY;HNRNPC (3)

HNRNPC (1)

HNRNPC (1)

HNRNPC (2)

HNRNPC (2)

HNRNPCL1 (1)

HNRNPCL1 (1)

HNRNPCL1 (2)

HNRNPCL1 (2)

RALYL (1)

RALYL (1)

RALYL (2)

RALYL (2)

Rb97D

Rb97D

RBFOX (1)

RBFOX (1)

RBFOX (2)

RBFOX (2)

RBFOX (3)

RBFOX (3)

RBFOX2

RBFOX2

RBFOX3

RBFOX3

RBM10 (1)

RBM10 (1)

RBM10 (2)

RBM10 (2)

RBM15B (4)

RBM15B (4)

RBM15B (1)

RBM15B (1)

RBM15B (2)

RBM15B (2)

RBM15B (3)

RBM15B (3)

RBM22

RBM22

RBM23 (1)

RBM23 (1)

RBM23 (2)

RBM23 (2)

RBM24/38 (1)

RBM24/38 (1)

RBM24/38 (2)

RBM24/38 (2)

RBM24/38 (3)

RBM24/38 (3)

RBM24 (1)

RBM24 (1)

RBM24 (2)

RBM24 (2)

RBM24 (3)

RBM24 (3)

RBM38

RBM38

sup-12

sup-12

RBM25 (1)

RBM25 (1)

RBM25 (2)

RBM25 (2)

RBM25 (3)

RBM25 (3)

RBM28

RBM28

RBM34

RBM34

RBM14 (1)

RBM14 (1)

RBM14 (2)

RBM14 (2)

RBM4 (1)

RBM4 (1)

RBM4 (2)

RBM4 (2)

RBM4 (3)

RBM4 (3)

RBM4 (4)

RBM4 (4)

RBM4 (5)

RBM4 (5)

RBM4B (1)

RBM4B (1)

RBM4B (2)

RBM4B (2)

RBM4B (3)

RBM4B (3)

RBM4B (4)

RBM4B (4)

RBM41 (1)

RBM41 (1)

RBM41 (2)

RBM41 (2)

RBM41 (3)

RBM41 (3)

RBM42 (1)

RBM42 (1)

RBM42 (2)

RBM42 (2)

RBM45 (1)

RBM45 (1)

RBM45 (2)

RBM45 (2)

RBM45 (3)

RBM45 (3)

RBM45 (4)

RBM45 (4)

RBM45 (5)

RBM45 (5)

RBM45 (6)

RBM45 (6)

RBM45 (7)

RBM45 (7)

RBM45 (8)

RBM45 (8)

RBM5 (1)

RBM5 (1)

RBM5 (2)

RBM5 (2)

RBM6 (1)

RBM6 (1)

RBM6 (2)

RBM6 (2)

RBM6 (3)

RBM6 (3)

RBM6 (4)

RBM6 (4)

RBM6 (5)

RBM6 (5)

RBM8A

RBM8A

RBMS1

RBMS1

RBMS2 (1)

RBMS2 (1)

RBMS2 (2)

RBMS2 (2)

RBMS2 (3)

RBMS2 (3)

RBMS3 (1)

RBMS3 (1)

RBMS3 (2)

RBMS3 (2)

RBMS3 (3)

RBMS3 (3)

RBMS3 (4)

RBMS3 (4)

RBMS3 (5)

RBMS3 (5)

RBMS3 (6)

RBMS3 (6)

RBPMS (1)

RBPMS (1)

RBPMS (2)

RBPMS (2)

RC3H1 (1)

RC3H1 (1)

RC3H1 (2)

RC3H1 (2)

RC3H1 (3)

RC3H1 (3)

RC3H1 (4)

RC3H1 (4)

Ref2

Ref2

rnp-2

rnp-2

Rnp4F

Rnp4F

RNPC3

RNPC3

Rox8

Rox8

Rsf1

Rsf1

rsp-1

rsp-1

rsp-2

rsp-2

SAMD4A

SAMD4A

SART3

SART3

SERBP1

SERBP1

SF1 (1)

SF1 (1)

SF1 (2)

SF1 (2)

SF1 (3)

SF1 (3)

SF1 (4)

SF1 (4)

SF3B4 (1)

SF3B4 (1)

SF3B4 (2)

SF3B4 (2)

SF3B6

SF3B6

sfa-1

sfa-1

shep (1)

shep (1)

shep (2)

shep (2)

shep (3)

shep (3)

sm

sm

SNRNP70

SNRNP70

SNRPA (1)

SNRPA (1)

SNRPA (2)

SNRPA (2)

SNRPA (3)

SNRPA (3)

SNRPA (4)

SNRPA (4)

sqd-1

sqd-1

SRSF1 (1)

SRSF1 (1)

SRSF1 (2)

SRSF1 (2)

SRSF1 (3)

SRSF1 (3)

SRSF1 (4)

SRSF1 (4)

SRSF1 (5)

SRSF1 (5)

SRSF1 (6)

SRSF1 (6)

SRSF9 (1)

SRSF9 (1)

SRSF9 (2)

SRSF9 (2)

SRSF9 (3)

SRSF9 (3)

SRSF9 (4)

SRSF9 (4)

SRSF9 (5)

SRSF9 (5)

srsf1a

srsf1a

SRSF10 (1)

SRSF10 (1)

SRSF10 (2)

SRSF10 (2)

SRSF10 (3)

SRSF10 (3)

SRSF10 (4)

SRSF10 (4)

SRSF10 (5)

SRSF10 (5)

SRSF10 (6)

SRSF10 (6)

SRSF11 (1)

SRSF11 (1)

SRSF11 (2)

SRSF11 (2)

SRSF2 (1)

SRSF2 (1)

SRSF2 (2)

SRSF2 (2)

SRSF2 (3)

SRSF2 (3)

SRSF2 (4)

SRSF2 (4)

Rbp1

Rbp1

Rbp1-like

Rbp1-like

SRSF3 (1)

SRSF3 (1)

SRSF3 (2)

SRSF3 (2)

SRSF4 (1)

SRSF4 (1)

SRSF4 (2)

SRSF4 (2)

SRSF4 (3)

SRSF4 (3)

SRSF6

SRSF6

SRSF5 (1)

SRSF5 (1)

SRSF5 (2)

SRSF5 (2)

SRSF5 (3)

SRSF5 (3)

SRSF5 (4)

SRSF5 (4)

SRSF5 (5)

SRSF5 (5)

SRSF8 (1)

SRSF8 (1)

SRSF8 (2)

SRSF8 (2)

ssx

ssx

sup-26

sup-26

HNRNPR (1)

HNRNPR (1)

HNRNPR (2)

HNRNPR (2)

HNRNPR (3)

HNRNPR (3)

Syp

Syp

T07F10.3

T07F10.3

TAF15 (1)

TAF15 (1)

TAF15 (2)

TAF15 (2)

TAF15 (3)

TAF15 (3)

TARDBP (3)

TARDBP (3)

TARDBP (4)

TARDBP (4)

TARDBP (1)

TARDBP (1)

TARDBP (2)

TARDBP (2)

tdp-1

tdp-1

TIA1 (1)

TIA1 (1)

TIA1 (2)

TIA1 (2)

TIA1 (3)

TIA1 (3)

TIA1 (4)

TIA1 (4)

TIA1 (5)

TIA1 (5)

TIA1 (6)

TIA1 (6)

tiar-1

tiar-1

tiar-3

tiar-3

tra2

tra2

TRA2A (1)

TRA2A (1)

TRA2A (2)

TRA2A (2)

TRIM56

TRIM56

TRIM71

TRIM71

TRNAU1AP (1)

TRNAU1AP (1)

TRNAU1AP (2)

TRNAU1AP (2)

TRNAU1AP (3)

TRNAU1AP (3)

TRNAU1AP (4)

TRNAU1AP (4)

TUT1

TUT1

U2AF2 (1)

U2AF2 (1)

U2AF2 (2)

U2AF2 (2)

U2AF2 (3)

U2AF2 (3)

U2AF2 (4)

U2AF2 (4)

UNK (1)

UNK (1)

UNK (2)

UNK (2)

UNK (3)

UNK (3)

UNK (4)

UNK (4)

wech

wech

Y57G11C.36

Y57G11C.36

Y59A8B.10

Y59A8B.10

YBX1 (1)

YBX1 (1)

YBX1 (2)

YBX1 (2)

YBX2

YBX2

ZC3H10

ZC3H10

ZC3H14

ZC3H14

CNBP

CNBP

ZCRB1 (1)

ZCRB1 (1)

ZCRB1 (2)

ZCRB1 (2)

ZFP36

ZFP36

ZNF326 (1)

ZNF326 (1)

ZNF326 (2)

ZNF326 (2)

ZNF326 (3)

ZNF326 (3)

ZNF326 (4)

ZNF326 (4)

ZNF638

ZNF638

To determine the age of motif hits across the 7 mammalian species, we identified the corresponding ancestral nodes for each of our species from the phylogeny derived from the Zoonomia consortium alignment of 447 mammalian genomes (Kuderna et al., Nature, 2023; Genereux et al., Nature, 2020; Cock et al., Biopython, 2009). We estimated the approximate age of the ancestral sequences from the TimeTree database (Kumar et al., MBE, 2022). We used halliftover to map motif-hit coordinates from the 7 mammals to their respective inferred ancestral sequences (Hickey et al., Bioinformatics, 2013). With bedtools, we then extended the lifted over coordinates by 6 bp on either side and extracted the corresponding ancestral sequences (Quinlan and Hall, Bioinformatics, 2010). We scanned the resulting sequences for their respective pre-calculated motif, and classified a sequence as a hit if its MSS score was at least in the 50th percentile. Finally, we estimated the approximate age of a motif hit by the age of the oldest ancestor for which the mapped region had a motif hit, with unmapped sequences classified as of Human origin.

The genomic region of motif instances identified by oRNAment at a threshold of 50% greatly correspond to the binding regions observed by eCLIP at a threshold of 3 fold change and a p-value of 0.001. Furthermore, the same number of random genomic regions as observed by oRNAment are seldom matching a binding region in eCLIP.

validation
Histogram showing the percent of motif instances observed by oRNAment and a corresponding number of random regions that are also observed as binding regions by eCLIP for 24 RBPs.

In oRNAment V2 we improved our database in the following aspects:

  • Motif instances are now predicted using the recent Ensembl release 115 (September 2025).
  • Ten new species have been added, bringing the total to 15 species covered in the database:
    • Human (Homo sapiens) Genome Assembly GRCh38
    • Mouse (Mus musculus) Genome Assembly GRCm39
    • Fruit Fly (Drosophila melanogaster) Genome Assembly BDGP6.54
    • Zebrafish (Danio rerio) Genome Assembly GRCz11
    • Roundworm (Caenorhabditis elegans) Genome Assembly WBcel235
    • Chimpanzee (Pan troglodytes) Genome Assembly Pan_tro_3.0
    • Rhesus macaque (Macaca mulatta) Genome Assembly Mmul_10
    • Olive baboon (Papio anubis) Genome Assembly Panubis1.0
    • Gray mouse lemur (Microcebus murinus) Genome Assembly Mmur_3.0
    • Rat (Rattus norvegicus) Genome Assembly GRCr8
    • Chicken (Gallus gallus) Genome Assembly bGalGal1.mat.broiler.GRCg7b
    • Western clawed frog (Xenopus tropicalis) Genome Assembly UCB_Xtro_10.0
    • Budding yeast (Saccharomyces cerevisiae) Genome Assembly R64-1-1
    • Lizard (Anolis carolinensis) Genome Assembly AnoCar2.0v2
    • Zebra finch (Taeniopygia guttata) Genome Assembly bTaeGut1_v1.p
  • Expanded RBP and motif coverage: the database now includes 2,089 RBPs across 15 species and 492 motifs. Homologous RBPs from different species are counted separately, though they typically share the same motif. The complete relationship table between the motifs and the RBPs can be found in the download page
  • Extended scan regions to include up to 150 nt of intronic sequence flanking exons, enabling coverage of splicing regulatory signals and other functions dependent on exon–intron boundaries. Intron are tagged through their host transcript, in the format of (transcript id)_(intron number)_(side), where side indicates whether it involves the 5' or 3' of the intron, or 5-3 if the entire intron is shorter than 300 nt.
  • Added search function to table views of motif instances to help navigate through all the motif instances identified in our database.
  • Added links to UCSC genome browser track hubs to interactively browse motif instances
  • Added estimated motif age for mammals to show the conservation of motifs

Please cite:

Louis Philip Benoit Bouvrette, Samantha Bovaird, Mathieu Blanchette, Eric Lécuyer oRNAment: A database of putative RNA binding protein target sites in the transcriptomes of model species. https://doi.org/10.1093/nar/gkz986

This site has been validated on Chrome, Firefox, Safari, and Vivaldi on both macOS and Windows environments. For best user experience we recommend Chrome or Vivaldi. Some functionality may not work if using Explorer on Windows.

The search by... forms

This tool allows you to search the database for a specific gene or transcript, for which you would like to know all putative RBP motif instances (by leaving the RBP choice as "All RBPs"), or for one RBP by choosing it from the drop down list. The choice of RBP also include the species to which it belongs, so the species-specific flag is ignored.

All choice fields are searchable.

In v2 only one gene is allowed at a time to reduce server pressure as by including intron and expanding our RBP list, a lot more motifs are detected than in the previous version.

Search_by_transcript_form

This tool allows you to search the database for a specific combination of attributes in a given species for which you would like to know all putative RBP motif instances.

Select the species and the biotype or region of interest. Notice that not all combinations of choices are compatible.

Search_by_Attributes_form

This tool allows you to search the database for a specific RBP from a specific species, at a specific similarity threshold between the PWM and the subsequence, for which you would like to know all its putative instances in all the coding and non-coding transcripts of the selected organism.

Select your organism and your RBP of interest. Input the threshold for motif score, whose range is limited between 0.5 and 1 to limit output table size while allowing for good overlap between motif prediction and in vitro binding essay results.

All choice fields are searchable.

Search_by_RPB_form

The results pages

The different search tools provide for various top-level graph analysis and a detailed table of all motif instances in each transcript. Kindly note that, by design, the database uses a "transcript-centric" view. Motifs instances coming from different isoform of a gene may have the same genomic coordinates. Therefore, a motif instance can be present more than once in the table, but will always represent two or more distinct transcripts of a gene.

The result page starts always with a call back of search criteria

call back

The following plots and tables are interactive and searchable when a lot of entries are expected.

The data for each plot and entire table of motif instances can be downloaded as file. Please note that the CSV may take a moment to generate.

When searching by gene name it could map to multiple Ensembl Gene ID if they are found on alternative contigs among other reasons. When searching by transcript, while we executed the scan on all transcripts available from the Ensembl database, due to varying nomenclature and versioning, it is possible that a specific gene name/ID will not be found in oRNAment.

Either case a table containing identified gene and transcripts is shown in the result page.

call back

The first plot shows the motif abundance through bar plot, histogram, tree map and searchable table. Different motifs of the same RBP are displayed individually.

motif abundance

The next plot shows the motif distribution across transcript regions. Please note that the entire squence of non-coding trancripts from a coding gene is classified as "Non Coding" despite it may share identical regions to UTRs from its coding counter parts. However since the "coding region" is technically undefined for these transcripts, it is not possible to distinguish UTR from the rest of the transcript.

region distributionback

Finally, all motif instances are shown as a table where it is possible to filter each column to facilitate browsing.

searchable, downloadable table

Please notice that not all combinations of search criteria will produce results. For example, choosing "snoRNA" for gene and transcript types while "CDS" for transcript region is not a valid combination, as snoRNAs do not have CDS. We are actively working to handle the selection logic to eleminate non-valid combinations.

The first plot shows the motif abundance through bar plot, histogram, tree map and searchable table. Different motifs of the same RBP are displayed individually.

motif abundance

Due to potential high row count of the motif instances, we do not offer browserable view of the table. A downloadable csv is still offered, with the total number of rows to help verifying the completeness of file.

download table

The first plot shows the biotypes of gene harboring the motif(s) of the RBP through bar plot, histogram, tree map and searchable table.

motif abundance

The next plot shows the motif distribution across transcript regions. Please note that the entire squence of non-coding trancripts from a coding gene is classified as "Non Coding" despite it may share identical regions to UTRs from its coding counter parts. However since the "coding region" is technically undefined for these transcripts, it is not possible to distinguish UTR from the rest of the transcript.

region distributionback

Due to potential high row count of the motif instances, we do not offer browserable view of the table. A downloadable csv is still offered, with the total number of rows to help verifying the completion of file.

download table

After the call back, all motifs from the same cluster (see above) are displayed with a marker shape assigned to each of them.

This is followed by the plots for transcripts and introns from the given gene or transcript. If a transcript or an intron does not contain any motifs from the motif cluster, the transcript or intron is omitted from this page.

On the left, the motifs are plotted with the transcript position as x-axis, and the motif score as y-axis, with the shade indicating its unpaired probability. On the right, we show the structure predicted by RNAfold and visualized by FornaContainer.

motif abundance

We decided to not predict structure for the intronic sequences, as it is unlikely to reflect the true structure of the intron in its context. In this case the unpaired probability is set to 1 solely for the purpose of display for introns.

motif abundance

The feature ID of an intron is derived from its host transcript, in the format of (transcript id)_(intron number)_(side), where side indicates whether it involves the 5' or 3' of the intron, or 5-3 if the entire intron is shorter than 300 nt. We do not include predicted structure for introns as the predictions based on 150 nt region are unlikely to reflect the true structure of introns within their context.