"The price of greatness is responsibility." Sir Winston Churchill


Search the IBPA



Top Menu

Menu Sidebar

IBPA Issues
About IBPA
IBPA Constitution
FAQ-s
IBPA Events
Individual Membership
Institutional Membership
IBPA Forums / Groups
Cooperation with IBPA
Links

Publications
IBPA Careers Newsletter
Past Issues
Industry Publications
Promote Yourself within the Industry
Submit Your Article

Career Center: Employers
Job Posting
Free Resume Database
Volunteers Database

Career Center: Job Seekers
Now Hiring
Submit Resume
Career Training
Nurses Careers in Biopharm
Scholarship Programs
Internship Programs
Resume Editing & Interview Coaching
Volunteer for the Industry
Download IBPA Career Info Brochure

Industry Directories and Listings
Pharmaceutical Companies
Contract Research Organizations
Professional Associations
Recruiters and Staffing Agencies
Clinical Research Centers
Consulting Companies
Education & Training Institutions
Jobs and Resume Searching Directories
Research and Development Companies
Industry Service Providers
List Your Company

Investor's Center
Offers
Calls

Contact IBPA
USAChapter
Canadian Chapter
European Chapter
Asian Chapter

Start Your Career in Biotech with IBPA Scholarship Programs
Untitled Document



Subscribe to our "Careers in the Biopharmaceutical Industry" newsletter:

Name*:

Email*:

City:

Country:

Phone:

To unsubscribe, click here

 

 

FASTA format

From Wikipedia, the free encyclopedia.

 

In bioinformatics, FASTA format is a file format used to exchange information between genetic sequence databases. Its format looks like this:

>SEQUENCE_1

;comment line 1 (optional)

MTEITAAMVKELRESTGAGMMDCKNALSETNGDFDKAVQLLREKGLGKAAKKADRLAAEG

LVSVKVSDDFTIAAMRPSYLSYEDLDMTFVENEYKALVAELEKENEERRRLKDPNKPEHK

IPQFASRKQLSDAILKEAEEKIKEELKAQGKPEKIWDNIIPGKMNSFIADNSQLDSKLTL

MGQFYVMDDKKTVEQVIAEKEKEFGGKIKIVEFICFEVGEGLEKKTEDFAAEVAAQL

>SEQUENCE_2

;comment line 1 (optional)

;comment line 2 (optional)

SATVSEINSETDFVAKNDQFIALTKDTTAHIQSNSLQSVEELHSSTINGVKFEEYLKSQI

ATIGENLVVRRFATLKAGANGVVNGYIHTNGRVGVVIAAACDSAEVASKSRDLLRQICMH

It consists of a header line (beginning with a '>') which gives a name and/or a unique identifier for the sequence, and often lots of other information too. Many different sequence databases use standarized headers, which helps when automatically extracting information from the header. Often the first 'word' of the header is a unique identifier for the sequence.

After the header line, one or more comments, distinguished by a semi-colon at the beginning of the line, may occur. Most databases and bioinformatics applications do not recognize these comments so their use is discouraged, but they are part of the official format.

After the header line and comments, one or more sequence lines may follow. Sequences may be protein sequences or DNA sequences, they must be shorther than 80 characters and can contain gaps or alignment characters (see sequence alignment).

FASTA format files often have file extensions like .fa, .mpfa, fna, or .fsa (and probably many more!).

The simple format of FASTA files makes them easy to manipulate using text processing tools and scripting languages like Perl.

The NCBI have gone so far as to define a standard for their fasta header (although generally this is a bit messy). The formatdb man page has this to say on the subject of FASTA format databases, "formatdb will automatically parse the SeqID and create indexes, but the database identifiers in the FASTA definition line must follow the conventions of the FASTA Defline Format."

However they do not give a difinitive description of the FASTA defline format, an attempt to create such a format is given below.

 GenBank                           gi|gi-number|gb|accession|locus

 EMBL Data Library                 gi|gi-number|emb|accession|locus

 DDBJ, DNA Database of Japan       gi|gi-number|dbj|accession|locus

 NBRF PIR                          pir||entry

 Protein Research Foundation       prf||name

 SWISS-PROT                        sp|accession|name

 Brookhaven Protein Data Bank (1)  pdb|entry|chain

 Brookhaven Protein Data Bank (2)  entry:chain|PDBID|CHAIN|SEQUENCE

 Patents                           pat|country|number 

 GenInfo Backbone Id               bbs|number 

 General database identifier       gnl|database|identifier

 NCBI Reference Sequence           ref|accession|locus

 Local Sequence identifier         lcl|identifier

[edit]

 

See also

[edit]


External links




Learn More About the Biopharmaceutical Industry and Clinical Research:


Category:

Logo sidebar
  • Analytical Chemistry
  • Bioinformatics
  • Biology
  • Biochemistry
  • Biotechnology
  • Biotechnology Companies
  • Cell Imaging
  • Chemistry
  • Chemists
  • Crystallography
  • Ecology
  • Environmentalism
  • Genetic Engineering
  • Genetically Modified Organisms
  • Genetics
  • Health
  • Health Care
  • Health Sciences
  • Medical Specialities
  • Medicine
  • Molecular Genetics
  • Pharmaceutical Industry
  • Pharmacy
  • Pharmacology

  • Powered by Wikipedia, the free encyclopedia. Articles were developed by IBPA volunteers.

    Logo sidebar

    A

    B

    C

    D

    E

    F

    G

    I

    K

    L

    M

    N

    P

    Q

    R

    S

    T


    Logo sidebar


    IBPA Sponsors and Active Supporters

    http://www.payoneer.com/
    Access Clinical Trials

    Access Clinical Trials
    Access Clinical Trials


    Allied Research International
    Allied Research International

    Altaspera Global Services Inc.
    Altaspera Global Services

    Financial Planning and Personal Insurance
    For Canadian Pharmaceutical Industry Executives


    Biorole Scientific Solutions
    Biorole Scientific Solutions

    CEREPROTEC INC. Development of Novel Neuroprotective Drugs
    CEREPROTEC INC. Development of Novel Neuroprotective Drugs

    Recruitment Advertising Agencies
    Recruitment Advertising Agencies

    Cellular Technology Ltd.
    Cellular Technology Ltd.

    Clinical Trial Network
    Free Database of Clinical Investigators

    ClinQua Clinical Trials Inc.
    ClinQua Clinical Trials Inc.

    Coronis Clinical Research Organization
    Coronis Clinical Research Organization

    CPIC Latin America
    CPIC Latin America

    Espoir Bridge Recruiters
    Espoir Bridge Recruiters

    Genentech
    Genentech

    ILS SA
    Independent Research and Laboratory Solutions

    Inova Health Research
    Inova Health Research, Inc.

    Kriger Research Group International
    Kriger Research Group International

    LCCT
    LCCT

    Metrics Research
    Complete Research Solutions on a Single Platform

    Pharmalef Developments
    Pharmalef Developments

    PrimeHealth Clinical Research Organization
    PrimeHealth Clinical Research Organization

    Research & Development RA SA
    Research & Development RA SA

    Scios Inc.
    Scios Inc. - Manufacturer of Health Care Products

    Scios Inc.
    Southeast Regional Research Group LLC.

    UniMR
    UniMR Clinical Research

    YM BioSciences
    YM BioSciences

    Become IBPA Sponsor
    Post Your Logo Here

    ©2004 International Biopharmaceutical Association Inc., all rights reserved
    Privacy Policy - Terms of Use

    Google