Amazon Comprehend Custom Entity Recognition: Setup Guide
Published on · Updated on
Corrected Comprehend training-data requirements and document-format support, replaced invalid IAM and endpoint examples, and clarified model evaluation, inference, and endpoint costs.
Amazon Comprehend custom entity recognition lets you train a model to find domain-specific entity types such as PART_ID, POLICY_NUMBER, or PRODUCT_CODE. It is useful when a phrase needs context to identify correctly or when the entity categories in your documents are not covered by Comprehend's built-in recognizer.
A custom recognizer predicts the entity types included in its training data. It does not automatically add built-in types such as PERSON or DATE, and a trained model is not a guaranteed match for every document. Evaluate predictions on examples from your own workload before relying on them. If you need noun phrases rather than labeled entities, see this guide to Amazon Comprehend key phrase extraction.
Prerequisites
Before training, choose an AWS Region where Amazon Comprehend is available and put the training data in an Amazon S3 bucket in that same Region. Keep the bucket private and organize input, training, and output objects under prefixes that you can scope in IAM.
Amazon Comprehend needs a data access role that trusts the Comprehend service and grants it access to the required S3 objects. Grant read access to training or job input data, and write access to the output prefix for asynchronous jobs. The identity that submits training or an asynchronous analysis job also needs permission to pass that role with iam:PassRole. Grant the application only the Comprehend actions and S3 access it needs; the service role and caller have separate permissions. For S3 policy concepts, see Amazon S3 access management.
For training, make sure you have realistic documents and consistent labels. For evaluation, prepare representative test documents that are separate from the examples used to build the model. The amount of data required depends on whether you use an entity list, plain-text annotations, or PDF annotations.
Prepare Your Data
Comprehend supports two training approaches. Both require documents that resemble the content you plan to analyze, but they provide different amounts of context to the model.
Entity lists for plain text
An entity list is a CSV of example strings and their custom type labels. You also provide a collection of unannotated documents containing those entities. Entity lists work only with plain-text documents; they are a simpler starting point when you can enumerate most of the relevant values. AWS recommends at least 25 entity matches for each type in the list and coverage of 80%–100% of positive mentions in the document corpus. Matches should occur in those training documents, and common words that could create false matches should be avoided.
Text,Type
P-1042,PART_ID
P-2088,PART_ID
AX-771,PART_ID
This is only a format example: a production list needs the recommended coverage, and its entries must appear in the unannotated document corpus. The entity-list method gives Comprehend less context than annotations, so it may produce more false positives when an entity string is ambiguous.
Annotations for context
Annotations mark each entity's exact location in source documents and associate it with a type. This lets the model learn from surrounding words, which is useful when similar text can represent different things in different contexts. Keep labels consistent, annotate every valid mention, and make offsets match the source text exactly. Avoid duplicate documents because they can contaminate the test data and distort metrics.
For plain-text annotations, AWS currently requires at least three annotated input documents and 25 annotations for each entity type. The CSV uses columns such as File, Line, Begin Offset, End Offset, and Type; line numbers start at zero. For example, in the line Replace pump P-1042 before restart., the span P-1042 begins at offset 13 and ends at offset 19.
If a plain-text annotation dataset has fewer than 50 annotations total and no separate test data, Comprehend reserves more than 10% of its documents for testing. Check the current annotation guidance for the behavior that applies to your data.
To analyze PDFs, Word documents, or images, train with annotated PDFs. AWS currently requires at least 250 input documents and 100 annotations per entity type for PDF annotation training. This model type supports English only. Once trained, it can analyze plain text, PDFs, Word documents, and JPG, PNG, or TIFF images. Text-only training does not enable those semi-structured inputs. Check the current annotation requirements and Comprehend quotas before assembling a dataset; these minimums make a training request eligible but do not guarantee model quality.
Create an Entity Recognizer
Decide which custom types to recognize and use clear uppercase labels with underscores, such as PART_ID or POLICY_NUMBER. A recognizer can include up to 25 custom entity types. Each model is trained for one language, and the training documents must use that language.
The following AWS CLI example starts training from an entity list and plain-text documents. Replace the example account, role, bucket, and Region with values from your account. The role must grant Amazon Comprehend access to the listed S3 training data.
aws comprehend create-entity-recognizer \
--recognizer-name part-id-recognizer \
--data-access-role-arn arn:aws:iam::123456789012:role/ComprehendDataAccessRole \
--input-data-config "EntityTypes=[{Type=PART_ID}],Documents={S3Uri=s3://example-bucket/train/documents/},EntityList={S3Uri=s3://example-bucket/train/entity-list.csv}" \
--language-code en \
--region us-east-1
For annotated text, configure the Annotations location instead of EntityList. For annotated PDF training, use the supported manifest and PDF annotation format described in the Amazon Comprehend custom entity recognition guide.
Train and Evaluate the Model
Training runs asynchronously. Use the returned recognizer ARN to check status:
aws comprehend describe-entity-recognizer \
--entity-recognizer-arn arn:aws:comprehend:us-east-1:123456789012:entity-recognizer/part-id-recognizer \
--region us-east-1
Wait until the recognizer status is TRAINED before using it. If the status is IN_ERROR, inspect the failure reason in the response and check the data format, S3 paths, Region, role trust policy, and permissions.
Review precision, recall, and F1 results for each entity and for the recognizer overall. If you provide a separate test set for annotation training, it must contain at least one annotation for each entity type. Entity-list training also supports a separate test document set, without an annotation CSV. When you do not supply test data, Comprehend reserves 10% of the input documents for testing; for plain-text annotation data with fewer than 50 total annotations, it reserves more than 10%. AWS describes these metrics as an estimate of model performance: they approximate results during API inference, rather than guarantee the same quality on every production document. The F1 scale is 0–100 for plain-text recognizers and 0–1 for PDF/Word recognizers, so check the model type before comparing values. See how Comprehend calculates custom recognizer metrics.
Use your errors to decide what to change. False positives may mean a list contains ambiguous terms or examples need more context. Missed entities may point to gaps in coverage, inconsistent boundaries, or a mismatch between training and production documents. Correct the underlying labels or data, train a new model, and compare it with a held-out set before replacing a working model.
Run Real-Time or Batch Inference
Use an endpoint when an application needs synchronous predictions. For larger S3 collections, use an asynchronous entity detection job. Both methods require a trained recognizer.
Real-time inference with an endpoint
Create an endpoint attached to the trained recognizer. The example ARN format below uses an entity-recognizer ARN returned by training:
aws comprehend create-endpoint \
--endpoint-name part-id-endpoint \
--model-arn arn:aws:comprehend:us-east-1:123456789012:entity-recognizer/part-id-recognizer \
--desired-inference-units 1 \
--region us-east-1
Endpoint creation takes time. Wait until its status is IN_SERVICE, then use the endpoint ARN from the create response with DetectEntities:
aws comprehend detect-entities \
--endpoint-arn arn:aws:comprehend:us-east-1:123456789012:entity-recognizer-endpoint/part-id-endpoint \
--language-code en \
--text "Replace pump P-1042 before restart." \
--region us-east-1
For a custom recognizer endpoint, Comprehend uses the model's language; the request's language code is ignored. A result includes detected text, custom type, offsets, and a confidence score. AWS sets per-inference-unit throughput limits of 100 characters per second and two documents per second. The endpoint continues to incur charges while it is active, so delete it when you no longer need real-time inference. See how Comprehend endpoints work and its current quotas for throughput and cost details.
Batch inference from Amazon S3
For an S3 collection, start an asynchronous job with the trained recognizer ARN, an input prefix, an output prefix, and a data access role. This example treats each input file as one document:
aws comprehend start-entities-detection-job \
--job-name part-id-batch-1 \
--entity-recognizer-arn arn:aws:comprehend:us-east-1:123456789012:entity-recognizer/part-id-recognizer \
--language-code en \
--input-data-config "S3Uri=s3://example-bucket/input/,InputFormat=ONE_DOC_PER_FILE" \
--output-data-config "S3Uri=s3://example-bucket/output/" \
--data-access-role-arn arn:aws:iam::123456789012:role/ComprehendDataAccessRole \
--region us-east-1
Use the returned job ID with describe-entities-detection-job to check progress. When the job completes, retrieve the generated output archive from the S3 prefix you supplied and inspect its results. A model trained with PDF annotations can process supported image, PDF, and Word inputs in analysis jobs; image and scanned-document text extraction requires the relevant Amazon Textract permissions and may incur Textract charges. Plain-text-trained recognizers accept text only.
CloudWatch can help track endpoint activity and throughput. To monitor recognition quality, sample predictions and compare them with human-verified labels; operational metrics alone do not show whether the extracted entities are correct.
Troubleshoot Common Problems
- Training cannot read an S3 object: Confirm the bucket is in the recognizer's Region, the object key and CSV headers are correct, and the data access role trusts Comprehend and grants access to the required prefix. Confirm that the caller can pass that role.
- A supported file is rejected: Verify that the recognizer was trained with PDF annotations for semi-structured input. A plain-text recognizer is limited to plain text; PDF-annotated recognizers support English only.
- Too many false positives: Review entity-list matches for common or ambiguous strings. Add representative context through annotations when a match's meaning depends on its surroundings.
- Expected spans are missed: Check annotation boundaries and labels against the original text, then review whether training examples cover the document styles and entity variants seen in production.
- An endpoint call fails: Confirm the endpoint belongs to an entity recognizer, its status is
IN_SERVICE, and the caller can invoke the required Comprehend operation.
Summary
Custom entity recognition in Amazon Comprehend can label domain-specific spans with types you define. Start with an entity list when you have a reliable list of values and plain-text inputs. Use annotations when context matters; use PDF annotations when you need to analyze supported PDF, Word, or image documents. Check current AWS data minimums, evaluate the recognizer on representative held-out documents, and choose between a billed real-time endpoint and asynchronous S3 jobs based on how your application needs results.