Abstract
Efforts to determine full-length cDNA sequences and characterize the cDNAs provide fundamental information to facilitate functional analysis of the transcripts, and recent studies have been extended to non-protein-coding transcripts (ncRNAs). New approaches are needed for
comprehensive analyses of massive amounts of cDNA sequences, which can be aimed not only at
protein-coding sequences (CDSs) but also at UTRs and ncRNAs. Batch Learning SOM (BL-SOM) is a powerful tool for extracting a wide range of genomic information. We constructed BL-SOM for pentanucleotide frequencies in ca. 30,000 full-length mouse cDNAs. In Fig.1, the 5' and 3' UTRs and CDSs of protein-coding cDNAs were separately analyzed, together with ncRNAs. Clear separation among these four functional categories (5' and 3' UTR, CDS, and ncRNA) was observed, and each functional category including ncRNA was divided into many sub-territories, that may reflect differences within each functional category. We next generated random sequences, which
maintained di- or tri-nucleotide frequency in individual ncRNAs; di- or tri-random, respectively, and mixed the random sequences with ncRNA sequences. In Fig.2, we constructed BL-SOM for tetra- or penta-nucleotide frequency in the ncRNA plus di- or tri-random sequences, respectively. Subdivision of ncRNA territory, that may reflect functional diversification, was again observed.