> For the complete documentation index, see [llms.txt](https://bigdata-2.gitbook.io/bd201notes/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://bigdata-2.gitbook.io/bd201notes/data_flow/nifi/streamsets.md).

# StreamSets

## Overview

* StreamSets Data Collector (SDC) is a lightweight, powerful design and execution engine that streams data in real time.
* SDC was released to the open source community in 2015.
* SDC is supported by StreamSets (founded 2014) that was 2018 Cloudera Partner Impact Awards Winner.

## StreamSets Data Collector Features

* Web-based user interface
* Highly configurable:
  * “At least once” and “At most once” delivery guarantees are supported.
  * Deep integration with the Hadoop ecosystem, including connectors for HDFS, HBase, Kafka and Solr.
  * Flexible deployment targets of pipelines to edge servers or to clusters.
  * Deployed as a Spark Streaming application or as a MapReduce job.
  * Embedded monitoring to provide runtime visibility to data flow performance.
* A key concept in SDC is the idea of pipeline

## What is pipeline?

* A pipeline describes the flow of data from the origin system to destination systems and defines how to transform the data along the way.
* A pipeline consists of a single origin stage to represent the origin system, multiple processor stages to transform data, and multiple destination stages to represent destination systems.

## Links

* [Introduction to StreamSets Data Collector](https://vimeo.com/141687653)
* [Building Data Pipelines with Apache Spark and StreamSets](https://streamsets.com/resources/video/building-data-pipelines-with-apache-spark-and-streamsets/)
* [A Short Course on the StreamSets Test](https://streamsets.com/resources/video/a-short-course-on-the-streamsets-test-framework/)
* [Getting Started with StreamSets Data Collector Edge](https://streamsets.com/resources/video/getting-started-with-streamsets-data-controller-edge/)
* [Ingesting Log Files into ElasticSearch Using StreamSets Data Collector](https://www.youtube.com/watch?v=bAIHFRf1vaw\&list=PLjInZV8YkDVQlpcS2yXkBnEI2xGas-0tT\&index=13\&t=0s)
* [RingCentral Scales Out with StreamSets](https://www.youtube.com/watch?time_continue=820\&v=KMs-1BgQB0Q)
* [Scripting processors and custom stages with StreamSets Data Collector](https://www.youtube.com/watch?v=hpwJyJhfddY\&list=PLMDP0ZREf8nl0Nk_5gTa9Rg8QtHzXt_Jw\&index=18)
* [Monitoring inventory changes with StreamSets Data Collector](https://www.youtube.com/watch?v=GtI6GRGL-iE\&list=PLMDP0ZREf8nl0Nk_5gTa9Rg8QtHzXt_Jw\&index=28)
* [Adaptive Data Cleansing with StreamSets and Cassandra](https://www.youtube.com/watch?v=DjRVEB2J7wo)
* [Ingest Data from Relational Databases to Cassandra with StreamSets](https://www.youtube.com/watch?v=V2ncOK4Uxz8)
* [Visualizing and Analyzing Salesforce Data with StreamSets and Neo4j](https://www.youtube.com/watch?v=AaqsNtIGN2s)
