dbifasta

Function

Description

dbifasta indexes a flat file database of one or more files, and builds EMBL CD-ROM format index files. This format is used by the software on the EMBL database CD-ROM distribution and by the Staden package in addition to EMBOSS, and appears to be the most generally used and publicly available index file format for these databases.

Usage

Command line arguments


Input file format

Output file format

dbifasta creates four index files. All are binary but with a simple format.

Data files

None.

Notes

The indexing method depends on each entry having a unique entry name. No allowance is made for two entries with the same name so it is not possible to index EMBL and EMBLNEW together.

Having created the EMBOSS indices for this file, a database can then be defined in the file emboss.defaults as something like:

DB emrod [
   type: N
   format: fasta
   method: emblcd
   directory: /data/embl/fasta
]  

Fields Indexed

By default, dbifasta will index the ID name and the accession number (if present).
If they are present in your database, you may also specify that dbifasta should index the Sequence Version and GI number and the words in the description by using the '-fields' qualifier with the appropriate values.

References

None.

Warnings

None.

Diagnostic Error Messages

Exit status It always exits with status 0.

Known bugs

None.

Author(s)

History

Target users

Comments