[{"data":1,"prerenderedAt":91},["ShallowReactive",2],{"content-guides-lambda-kafka-using-kafka":3},{"markdown":4,"frontMatterAttributes":5,"bodyRaw":8,"contributorPaths":9,"menuType":11,"contentName":12,"slug":13,"readingTime":14,"menu":15,"nextDocItem":33,"previousDocItem":19},"\u003Cp>\u003Ca href=\"https:\u002F\u002Fkafka.apache.org\u002F\">Apache Kafka\u003C\u002Fa> is a popular open-source platform for building real-time streaming data pipelines and applications. More than \u003Ca href=\"https:\u002F\u002Fkafka.apache.org\u002F\">80% of all Fortune 100 companies\u003C\u002Fa> use Kafka to modernize their data architecture.\u003C\u002Fp>\n\u003Cp>Kafka has several use-cases:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Real-time web and log analytics\u003C\u002Fli>\n\u003Cli>Transaction and event sourcing\u003C\u002Fli>\n\u003Cli>Messaging\u003C\u002Fli>\n\u003Cli>Decoupled microservices\u003C\u002Fli>\n\u003Cli>Event-Driven-Architectures\u003C\u002Fli>\n\u003Cli>Streaming ETL\u003C\u002Fli>\n\u003Cli>Change data capture\u003C\u002Fli>\n\u003Cli>Metrics and log aggregation\u003C\u002Fli>\n\u003Cli>Streaming ML\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch3>Designing a Kafka streaming application\u003C\u002Fh3>\n\u003Cp>There are a number of different ways to run Kafka.\u003C\u002Fp>\n\u003Cp>You can deploy and manage your own Kafka solution on-premises or in the cloud on \u003Ca href=\"https:\u002F\u002Faws.amazon.com\u002Fec2\u002F\">Amazon EC2\u003C\u002Fa>. For more information on hosting Kafka yourself, read \u003Ca href=\"http:\u002F\u002Faws.amazon.com\u002Fblogs\u002Fbig-data\u002Fbest-practices-for-running-apache-kafka-on-aws\u002F\">Best Practices for Running Apache Kafka on AWS\u003C\u002Fa> on the AWS Big Data Blog.\u003C\u002Fp>\n\u003Cp>\u003Ca href=\"https:\u002F\u002Faws.amazon.com\u002Fmsk\u002F\">Amazon Managed Streaming for Apache Kafka (MSK)\u003C\u002Fa> is a fully managed service that makes it easier for you to build and run applications that use Kafka to process streaming data. You can populate data lakes, stream changes to and from databases, and power machine learning and analytics applications.\u003C\u002Fp>\n\u003Cp>Amazon MSK has the following features:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Automate provisioning, configuring, and tuning\u003C\u002Fstrong>. Reduce operational overhead, including the provisioning, configuration, and maintenance of Apache Kafka and Kafka Connect clusters. You can configure your application for high availability across multiple Availability Zones.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Fully compatible with open-source Apache Kafka\u003C\u002Fstrong>. Use applications and tools built for Apache Kafka out of the box, including \u003Ca href=\"https:\u002F\u002Fcwiki.apache.org\u002Fconfluence\u002Fpages\u002Fviewpage.action?pageId=27846330\">MirrorMaker\u003C\u002Fa>, \u003Ca href=\"https:\u002F\u002Fflink.apache.org\u002F\">Apache Flink\u003C\u002Fa>, and \u003Ca href=\"https:\u002F\u002Fprometheus.io\u002F\">Prometheus\u003C\u002Fa>. Use the native Apache Kafka APIs, without having to change application code. You can use multiple open-source versions of Kafka.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Highly secure\u003C\u002Fstrong>. Easily deploy secure, production-ready applications using native integrations to an Amazon Virtual Private Cloud (VPC). You can use \u003Ca href=\"https:\u002F\u002Faws.amazon.com\u002Fiam\u002F\">AWS Identity and Access Management\u003C\u002Fa> (IAM) for simple authentication and authorization.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Lower cost\u003C\u002Fstrong>. Keep costs low with fully managed Apache Kafka, offered at as low as 1\u002F13th the cost of other providers. Data replication between Availability Zones is included at no additional cost with MSK.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>\u003Ca href=\"https:\u002F\u002Fdocs.aws.amazon.com\u002Fmsk\u002Flatest\u002Fdeveloperguide\u002Fserverless.html\">MSK Serverless\u003C\u002Fa> is a cluster type for Amazon MSK that allows you to run Kafka without having to manage and scale cluster capacity. It automatically provisions and scales capacity while managing the partitions in your topic, so you can stream data without thinking about right-sizing or scaling clusters. MSK Serverless offers a throughput-based pricing model, so you pay only for what you use. Consider using a serverless cluster if your applications need on-demand streaming capacity that scales up and down automatically.\u003C\u002Fp>\n\u003Cp>\u003Ca href=\"https:\u002F\u002Fdocs.aws.amazon.com\u002Fmsk\u002Flatest\u002Fdeveloperguide\u002Fmsk-broker-types-express.html\">MSK Express Brokers\u003C\u002Fa> for MSK Provisioned make Apache Kafka simpler to manage, more cost-effective to run at scale, and more elastic with the low latency you expect. Brokers include pay-as-you-go storage that scales automatically and requires no sizing, provisioning, or proactive monitoring. Depending on the instance size selected, each broker node can provide up to 3x more throughput per broker, scale up to 20x faster, and recover 90% quicker compared to standard Apache Kafka brokers. Express brokers come pre-configured with Amazon MSK’s best practice defaults and enforce client throughput quotas to minimize resource contention between clients and Kafka’s background operations. \u003C\u002Fp>\n\u003Cp>You can use \u003Ca href=\"https:\u002F\u002Fwww.confluent.io\u002Flp\u002Fconfluent-cloud\">Confluent Cloud\u003C\u002Fa> as a fully managed Kafka data streaming platform or use \u003Ca href=\"https:\u002F\u002Fwww.confluent.io\u002Fproduct\u002Fconfluent-platform\u002F\">Confluent Platform\u003C\u002Fa> if you prefer to deploy and manage the platform yourself. You can use \u003Ca href=\"https:\u002F\u002Fdocs.confluent.io\u002Fplatform\u002Fcurrent\u002Fconnect\u002Findex.html\">Kafka Connect\u003C\u002Fa> or \u003Ca href=\"https:\u002F\u002Fdocs.confluent.io\u002Fcloud\u002Fcurrent\u002Fconnectors\u002Findex.html\">Confluent Cloud Connectors\u003C\u002Fa> to send data to and from a number of AWS sources and targets (sink connectors).\u003C\u002Fp>\n\u003Ch3>Kafka architectural patterns\u003C\u002Fh3>\n\u003Ch4>Clusters, topics, and records\u003C\u002Fh4>\n\u003Cp>Kafka clusters store a stream of related records in an ordered, immutable log called a topic. Kafka records are also called messages or events. A Kafka record includes a key and a value, along with headers and a timestamp.\u003C\u002Fp>\n\u003Ch4>Producers and consumers\u003C\u002Fh4>\n\u003Cimg src=\"\u002Fassets\u002Fexternal\u002Fguides\u002Flambda-kafka\u002Fassets\u002Fimages\u002Fkafka-lambda-producers-consumers.png\">\n\n\u003Cp>Producers create records and send them to one or many topics on a Kafka cluster.\u003C\u002Fp>\n\u003Cp>Consumers read and process streams of records. The record consumer chooses where to start reading the stream. This can be from the earliest record available on the stream or somewhere in the middle. The consumer continues to read and process records until new records arrive. You can also choose to process only new records, ignoring records already in the stream.\u003C\u002Fp>\n\u003Cp>Record streams are persistent in Kafka as Kafka acts as a record store. Kafka does not remove records once consumers process them. You rather configure record retention periods. When the record retention period expires, Kafka removes the records.\u003C\u002Fp>\n\u003Ch4>Partitions\u003C\u002Fh4>\n\u003Cp>Topics can be split into partitions for improved scalability and throughput. Each partition is a single log file where Kafka writes records in an append-only fashion.\u003C\u002Fp>\n\u003Cimg src=\"\u002Fassets\u002Fexternal\u002Fguides\u002Flambda-kafka\u002Fassets\u002Fimages\u002Fkafka-lambda-partitions.png\">\n\n\u003Cp>When producers send records to a stream, the partition key determines which partition it routes to. Kafka hashes the partition key and uses the result to map the record to a specific partition. Messages with the same partition key route to the same partition. If producers don&#39;t specify a partition key, records are distributed round-robin across all the topic&#39;s partitions. If the topic has a single partition, the partition key has no effect, all records route to the same partition. If the topic has multiple partitions, Kafka assigns records to partitions by hashing the partition key and mapping to individual partitions.\u003C\u002Fp>\n\u003Cimg src=\"\u002Fassets\u002Fexternal\u002Fguides\u002Flambda-kafka\u002Fassets\u002Fimages\u002Fkafka-lambda-partition-key.png\">\n\n\u003Cp>Records with the same key always write to the same partition, and records in a partition are always in order.\u003C\u002Fp>\n\u003Cp>You can influence the partition assignment depending on your choice of partition key.\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Random:\u003C\u002Fstrong> A random value results in random hash, so records are randomly sent to different partitions. This effectively load balances records across all available partitions.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Time-based:\u003C\u002Fstrong> A timestamp value may cause groups of records to be sent to a single partition, if the records arrive at the same time. The identical timestamp results in an identical hash.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Application-specific:\u003C\u002Fstrong> If you produce records that are all related to a particular customer, you can use the \u003Cem>customerID\u003C\u002Fem> as the key. All records for that customer are routed to the same partition and always arrive in order. This can be useful for downstream aggregation logic but may limit the capacity of records per \u003Cem>customerID\u003C\u002Fem>.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch4>Offsets\u003C\u002Fh4>\n\u003Cp>Records in partitions are each assigned a sequential identifier called the offset. This is unique for each record within the partition and incrementally tracks which record a particular consumer is processing.\u003C\u002Fp>\n\u003Cimg src=\"\u002Fassets\u002Fexternal\u002Fguides\u002Flambda-kafka\u002Fassets\u002Fimages\u002Fkafka-lambda-consumer-offset.png\">\n",{"title":6,"order":7},"Using Kafka to build your streaming application",2,"[Apache Kafka](https:\u002F\u002Fkafka.apache.org\u002F) is a popular open-source platform for building real-time streaming data pipelines and applications. More than [80% of all Fortune 100 companies](https:\u002F\u002Fkafka.apache.org\u002F) use Kafka to modernize their data architecture.\n\nKafka has several use-cases:\n\n- Real-time web and log analytics\n- Transaction and event sourcing\n- Messaging\n- Decoupled microservices\n- Event-Driven-Architectures\n- Streaming ETL\n- Change data capture\n- Metrics and log aggregation\n- Streaming ML\n\n### Designing a Kafka streaming application\n\nThere are a number of different ways to run Kafka.\n\nYou can deploy and manage your own Kafka solution on-premises or in the cloud on [Amazon EC2](https:\u002F\u002Faws.amazon.com\u002Fec2\u002F). For more information on hosting Kafka yourself, read [Best Practices for Running Apache Kafka on AWS](http:\u002F\u002Faws.amazon.com\u002Fblogs\u002Fbig-data\u002Fbest-practices-for-running-apache-kafka-on-aws\u002F) on the AWS Big Data Blog.\n\n[Amazon Managed Streaming for Apache Kafka (MSK)](https:\u002F\u002Faws.amazon.com\u002Fmsk\u002F) is a fully managed service that makes it easier for you to build and run applications that use Kafka to process streaming data. You can populate data lakes, stream changes to and from databases, and power machine learning and analytics applications.\n\nAmazon MSK has the following features:\n\n- **Automate provisioning, configuring, and tuning**. Reduce operational overhead, including the provisioning, configuration, and maintenance of Apache Kafka and Kafka Connect clusters. You can configure your application for high availability across multiple Availability Zones.\n- **Fully compatible with open-source Apache Kafka**. Use applications and tools built for Apache Kafka out of the box, including [MirrorMaker](https:\u002F\u002Fcwiki.apache.org\u002Fconfluence\u002Fpages\u002Fviewpage.action?pageId=27846330), [Apache Flink](https:\u002F\u002Fflink.apache.org\u002F), and [Prometheus](https:\u002F\u002Fprometheus.io\u002F). Use the native Apache Kafka APIs, without having to change application code. You can use multiple open-source versions of Kafka.\n- **Highly secure**. Easily deploy secure, production-ready applications using native integrations to an Amazon Virtual Private Cloud (VPC). You can use [AWS Identity and Access Management](https:\u002F\u002Faws.amazon.com\u002Fiam\u002F) (IAM) for simple authentication and authorization.\n- **Lower cost**. Keep costs low with fully managed Apache Kafka, offered at as low as 1\u002F13th the cost of other providers. Data replication between Availability Zones is included at no additional cost with MSK.\n\n[MSK Serverless](https:\u002F\u002Fdocs.aws.amazon.com\u002Fmsk\u002Flatest\u002Fdeveloperguide\u002Fserverless.html) is a cluster type for Amazon MSK that allows you to run Kafka without having to manage and scale cluster capacity. It automatically provisions and scales capacity while managing the partitions in your topic, so you can stream data without thinking about right-sizing or scaling clusters. MSK Serverless offers a throughput-based pricing model, so you pay only for what you use. Consider using a serverless cluster if your applications need on-demand streaming capacity that scales up and down automatically.\n\n[MSK Express Brokers](https:\u002F\u002Fdocs.aws.amazon.com\u002Fmsk\u002Flatest\u002Fdeveloperguide\u002Fmsk-broker-types-express.html) for MSK Provisioned make Apache Kafka simpler to manage, more cost-effective to run at scale, and more elastic with the low latency you expect. Brokers include pay-as-you-go storage that scales automatically and requires no sizing, provisioning, or proactive monitoring. Depending on the instance size selected, each broker node can provide up to 3x more throughput per broker, scale up to 20x faster, and recover 90% quicker compared to standard Apache Kafka brokers. Express brokers come pre-configured with Amazon MSK’s best practice defaults and enforce client throughput quotas to minimize resource contention between clients and Kafka’s background operations. \n\nYou can use [Confluent Cloud](https:\u002F\u002Fwww.confluent.io\u002Flp\u002Fconfluent-cloud) as a fully managed Kafka data streaming platform or use [Confluent Platform](https:\u002F\u002Fwww.confluent.io\u002Fproduct\u002Fconfluent-platform\u002F) if you prefer to deploy and manage the platform yourself. You can use [Kafka Connect](https:\u002F\u002Fdocs.confluent.io\u002Fplatform\u002Fcurrent\u002Fconnect\u002Findex.html) or [Confluent Cloud Connectors](https:\u002F\u002Fdocs.confluent.io\u002Fcloud\u002Fcurrent\u002Fconnectors\u002Findex.html) to send data to and from a number of AWS sources and targets (sink connectors).\n\n### Kafka architectural patterns\n\n#### Clusters, topics, and records\n\nKafka clusters store a stream of related records in an ordered, immutable log called a topic. Kafka records are also called messages or events. A Kafka record includes a key and a value, along with headers and a timestamp.\n\n#### Producers and consumers\n\n\u003Cimg src=\"\u002Fassets\u002Fexternal\u002Fguides\u002Flambda-kafka\u002Fassets\u002Fimages\u002Fkafka-lambda-producers-consumers.png\">\n\nProducers create records and send them to one or many topics on a Kafka cluster.\n\nConsumers read and process streams of records. The record consumer chooses where to start reading the stream. This can be from the earliest record available on the stream or somewhere in the middle. The consumer continues to read and process records until new records arrive. You can also choose to process only new records, ignoring records already in the stream.\n\nRecord streams are persistent in Kafka as Kafka acts as a record store. Kafka does not remove records once consumers process them. You rather configure record retention periods. When the record retention period expires, Kafka removes the records.\n\n#### Partitions\n\nTopics can be split into partitions for improved scalability and throughput. Each partition is a single log file where Kafka writes records in an append-only fashion.\n\n\u003Cimg src=\"\u002Fassets\u002Fexternal\u002Fguides\u002Flambda-kafka\u002Fassets\u002Fimages\u002Fkafka-lambda-partitions.png\">\n\nWhen producers send records to a stream, the partition key determines which partition it routes to. Kafka hashes the partition key and uses the result to map the record to a specific partition. Messages with the same partition key route to the same partition. If producers don't specify a partition key, records are distributed round-robin across all the topic's partitions. If the topic has a single partition, the partition key has no effect, all records route to the same partition. If the topic has multiple partitions, Kafka assigns records to partitions by hashing the partition key and mapping to individual partitions.\n\n\u003Cimg src=\"\u002Fassets\u002Fexternal\u002Fguides\u002Flambda-kafka\u002Fassets\u002Fimages\u002Fkafka-lambda-partition-key.png\">\n\nRecords with the same key always write to the same partition, and records in a partition are always in order.\n\nYou can influence the partition assignment depending on your choice of partition key.\n\n- **Random:** A random value results in random hash, so records are randomly sent to different partitions. This effectively load balances records across all available partitions.\n- **Time-based:** A timestamp value may cause groups of records to be sent to a single partition, if the records arrive at the same time. The identical timestamp results in an identical hash.\n- **Application-specific:** If you produce records that are all related to a particular customer, you can use the _customerID_ as the key. All records for that customer are routed to the same partition and always arrive in order. This can be useful for downstream aggregation logic but may limit the capacity of records per _customerID_.\n\n#### Offsets\n\nRecords in partitions are each assigned a sequential identifier called the offset. This is unique for each record within the partition and incrementally tracks which record a particular consumer is processing.\n\n\u003Cimg src=\"\u002Fassets\u002Fexternal\u002Fguides\u002Flambda-kafka\u002Fassets\u002Fimages\u002Fkafka-lambda-consumer-offset.png\">\n",[10],"content\u002Fcontributors\u002Fjulian-wood.json","LESSON","Using AWS Lambda to process Apache Kafka streams","\u002Fcontent\u002Fguides\u002Flambda-kafka\u002Fusing-kafka","6 min",[16,25,30,38,47,55,64,73,82],{"title":17,"content":18},"Introducing streaming",[19],{"title":17,"order":20,"time":21,"path":22,"id":23,"link":24},1,"4 min","introduction","introduction.md","\u002Fcontent\u002Fguides\u002Flambda-kafka\u002Fintroduction",{"title":6,"content":26},[27],{"title":6,"order":7,"time":14,"path":28,"id":29,"link":13},"using-kafka","using-kafka.md",{"title":31,"content":32},"Using Lambda to consume records from Kafka",[33],{"title":31,"order":34,"time":21,"path":35,"id":36,"link":37},3,"using-lambda","using-lambda.md","\u002Fcontent\u002Fguides\u002Flambda-kafka\u002Fusing-lambda",{"title":39,"content":40},"Configuring Kafka and Lambda",[41],{"title":39,"order":42,"time":43,"path":44,"id":45,"link":46},4,"9 min","configuring-kafka-lambda","configuring-kafka-lambda.md","\u002Fcontent\u002Fguides\u002Flambda-kafka\u002Fconfiguring-kafka-lambda",{"title":48,"content":49},"Processing Kafka streams using the Lambda ESM",[50],{"title":48,"order":51,"time":43,"path":52,"id":53,"link":54},5,"processing-esm","processing-esm.md","\u002Fcontent\u002Fguides\u002Flambda-kafka\u002Fprocessing-esm",{"title":56,"content":57},"Scaling and throughput",[58],{"title":56,"order":59,"time":60,"path":61,"id":62,"link":63},6,"5 min","scaling-throughput","scaling-throughput.md","\u002Fcontent\u002Fguides\u002Flambda-kafka\u002Fscaling-throughput",{"title":65,"content":66},"Monitoring and observability",[67],{"title":65,"order":68,"time":69,"path":70,"id":71,"link":72},7,"2 min","monitoring-observability","monitoring-observability.md","\u002Fcontent\u002Fguides\u002Flambda-kafka\u002Fmonitoring-observability",{"title":74,"content":75},"Troubleshooting",[76],{"title":74,"order":77,"time":78,"path":79,"id":80,"link":81},8,"3 min","troubleshooting","troubleshooting.md","\u002Fcontent\u002Fguides\u002Flambda-kafka\u002Ftroubleshooting",{"title":83,"content":84},"Resources",[85],{"title":83,"order":86,"time":87,"path":88,"id":89,"link":90},9,"1 min","resources","resources.md","\u002Fcontent\u002Fguides\u002Flambda-kafka\u002Fresources",1789814105348]