The public API requires no login, API key, or browser cookies. It searches the same MAG database as the web query form. Save the returned job_id to retrieve your results later.

Base URL: https://pzlast.nig.ac.jp/pzlast/mag

Check the JSON response, not only the HTTP status. For compatibility with PZLAST, application errors are returned as HTTP 200 with response.Status = "Error" and an explanation in response.Message. Download endpoints return JSON on errors instead of a result file. The examples below check both kinds of response.

Quick start

These examples use Bash, curl, and jq. Copy the three shell blocks into one script and run it with bash. Replace ./my_sequence.fasta with a single- or multi-FASTA file containing protein sequences.

1. Submit a search

Upload the FASTA file as multipart form data named fasta_file. Search options belong in the URL query string. This example requests up to 10 hits per query sequence.

#!/usr/bin/env bash
set -euo pipefail

BASE="https://pzlast.nig.ac.jp/pzlast/mag"
reply=$(curl --fail --silent --show-error \
  --connect-timeout 10 --max-time 120 \
  --form 'fasta_file=@./my_sequence.fasta' \
  "$BASE/job/submit?max_out=10&e_value=1e-8&q_stride=1")

job_id=$(printf '%s' "$reply" | jq -er '
  .response | if .Status == "Success" then
    .job_id | select(type == "string" and length > 0)
  else error(.Message // "Submission failed") end')
printf 'Job ID: %s\n' "$job_id"

A successful submission returns:

{
  "response": {
    "Status": "Success",
    "Message": "Success. Record inserted.",
    "job_id": "YOUR_JOB_ID"
  }
}

2. Wait for results

Poll the status endpoint every 10 seconds. Only COMPLETE means downloadable results are ready; DONE still requires result conversion. This loop stops on an error or after a one-hour polling window (plus any in-flight request). Increase the window if necessary, or resume later with the saved job ID.

deadline=$((SECONDS + 3600))
ready=false
while (( SECONDS < deadline )); do
  reply=$(curl --fail --silent --show-error \
    --connect-timeout 10 --max-time 60 \
    --get --data-urlencode "job_id=$job_id" \
    "$BASE/job/status")
  state=$(printf '%s' "$reply" | jq -er '
    .response | if .Status == "Success" then .job_status
    else error(.Message // "Status request failed") end')
  printf 'Status: %s\n' "$state"
  case "$state" in
    COMPLETE) ready=true; break ;;
    ERROR) printf 'Search failed. Response: %s\n' "$reply" >&2; exit 1 ;;
    WAITING|RUNNING|DONE|CONVERTING) sleep 10 ;;
    *) printf 'Unexpected response: %s\n' "$reply" >&2; exit 1 ;;
  esac
done
if [ "$ready" != true ]; then
  printf 'Polling timed out. Keep job ID %s and check again later.\n' "$job_id" >&2
  exit 1
fi

3. Download the results

The helper checks the response content type before saving a file, so an API error cannot silently become a CSV or FASTA file. An empty result file is valid when a completed search has no hits.

download_result() {
  local endpoint="$1" output="$2" temporary content_type
  temporary=$(mktemp)
  if ! content_type=$(curl --fail --silent --show-error \
    --connect-timeout 10 --max-time 300 \
    --get --data-urlencode "job_id=$job_id" \
    --output "$temporary" --write-out '%{content_type}' \
    "$BASE/result/$endpoint"); then
    rm -f "$temporary"
    return 1
  fi
  case "$content_type" in
    application/json*)
      cat "$temporary" >&2
      rm -f "$temporary"
      return 1 ;;
  esac
  mv "$temporary" "$output"
}

download_result get_csv "${job_id}.csv"
download_result get_fasta "${job_id}.fasta"
download_result get_genome_csv "${job_id}_genomes.csv"

To inspect a saved job in the browser, enter its ID on the Result page. API requests do not require the browser's History list.

Submission parameters

POST /pzlast/mag/job/submit

fasta_file is required in multipart/form-data. Use protein sequences, not nucleotide sequences. The following URL query parameters are optional.

ParameterDefaultMeaning
max_out1Positive integer: maximum hits returned per query sequence. Query count × max_out must not exceed 100000.
e_value1e-8Finite positive number: E-value cutoff. Smaller values require stronger sequence similarity.
q_stride1Positive integer: query stride for the search engine. Keep the default for the normal search setting.
other_paramsEmptyOptional legacy parameter string, up to 255 characters. Normally omit this parameter.

Parameters from the earlier PZLAST service such as search_mode, target_meo, and target_sample do not apply to PZLAST-MAG. Searches use the MAG database without an MEO or sample filter.

Job status

GET /pzlast/mag/job/status?job_id=YOUR_JOB_ID

This endpoint reads job progress without changing it. job_id is required. A successful status request returns response.Status = "Success", even when the search itself has failed; inspect job_status as well.

{
  "response": {
    "Status": "Success",
    "job_id": "YOUR_JOB_ID",
    "job_status": "COMPLETE",
    "ready": true,
    "under_maintenance": false
  }
}
job_statusMeaningNext action
WAITINGQueued for computation.Wait and poll again.
RUNNINGSequence search is in progress.Wait and poll again.
DONEComputation finished; results await conversion.Wait and poll again.
CONVERTINGDownloadable results are being prepared.Wait and poll again.
COMPLETEResults are ready, including searches with no hits.Download results; ready is true.
ERRORThe search failed.Stop polling and try submitting again later.

ready is false for every status other than COMPLETE. under_maintenance reports the service maintenance setting. Use the status endpoint above for polling; /mag/job/get is an internal worker endpoint that assigns a job and changes its status.

Result formats

All result endpoints use GET and require job_id as a URL query parameter. Wait for COMPLETE before retrieving results.

EndpointOutput
/pzlast/mag/result/get_csvCSV table of individual protein sequence hits.
/pzlast/mag/result/get_fastaFASTA file of hit reference protein sequences.
/pzlast/mag/result/get_genome_csvCSV table aggregating hits by MAG, including query completion.
/pzlast/mag/result/get_recordsJSON with a response array of individual hit records.
/pzlast/mag/result/get_completionJSON with a response array of MAG completion records.
/pzlast/mag/result/get_binRaw binary search output for advanced use. Optional ext: fixbin (default), hdrbin, bdybin, or sidbin.

For example, inspect individual hits as JSON:

curl --fail --silent --show-error \
  --get --data-urlencode "job_id=$job_id" \
  "$BASE/result/get_records" | jq .

The JSON record endpoints return arrays under response when hits exist. With no hits, they return the success message below instead. CSV and FASTA downloads are empty files for a completed search with no hits.

{
  "response": {
    "Status": "Success",
    "Message": "No hit. Your job has been done but no hits found."
  }
}

Limits and errors

  • Each protein sequence must contain 10 to 10000 amino acids.
  • Submit at most 10000 sequences and 100000 total amino acids per job.
  • The number of query sequences multiplied by max_out must be at most 100000.
  • At most 10 jobs may be waiting or running across the service. If the queue is full, wait before submitting again.
  • During maintenance, submitted jobs remain queued until processing resumes.
  • Jobs and result files are eligible for removal 14 days after the job was last updated. Download and keep the files you need.

Missing files, invalid FASTA or options, exceeded limits, unknown or expired job IDs, unfinished searches, and failed searches return an error description in JSON. For example:

{
  "response": {
    "Status": "Error",
    "Message": "Error. No job_id specified."
  }
}

curl --fail detects HTTP or transport failures, but it does not detect these HTTP 200 API errors. Always inspect response.Status for object responses and the content type of downloads. If a submission response is lost, its job may already have been registered; avoid automatically repeating the POST request.