---
title: Business intelligence from your VPC Flow Logs!
description: Go from raw VPC Flow Log data to business intelligence in just a few minutes.
---

Revinate leaves their ELK stack behind to find huge gains with ChaosSearch -- Read More!

[Revinate leaves their ELK stack behind to find huge gains with ChaosSearch -- Read More!

](https://www.chaossearch.io/resources/customer-stories/revinate)

[![ChaosSearch](https://www.chaossearch.io/hubfs/2021%20Website/logo.svg) ](https://www.chaossearch.io/)

![Gartner Cool Vendor 2023](https://www.chaossearch.io/hubfs/C2020/Logos/Gartner%20Cool%20Vendor%202023.png)

[![Start Free Trial](https://no-cache.hubspot.com/cta/default/4020721/a1eeee12-14a2-41d7-a38b-11e32980b916.png)](https://cta-redirect.hubspot.com/cta/redirect/4020721/a1eeee12-14a2-41d7-a38b-11e32980b916)

# Log Analytics Unleashed

## Business intelligence from your VPC Flow Logs!

Go from raw VPC Flow Log data to business intelligence in just a few minutes. Troubleshoot potential security issues, ensure that network access rules work as expected, and much more!

In this webcast we will cover:

- Quickly uncover unusual security activities and where a connection originated from
- Define monitors to proactively notify on protocols and port numbers used for requests
- Discover what roadblocks you may face for rejected traffic
- Leverage ChaosSearch to Know Better® about VPC flow traffic

 

 

## Interested in scheduling a brief intro call to see how ChaosSearch can accelerate your analytics?

Yes - show me the calendar!

## Future-Proof Your Analytics at Scale

[![Get a Demo](https://no-cache.hubspot.com/cta/default/4020721/0ae4b1b0-9ec0-4169-80cc-64fda5fd56db.png)](https://cta-redirect.hubspot.com/cta/redirect/4020721/0ae4b1b0-9ec0-4169-80cc-64fda5fd56db)

©2024, ChaosSearch®, Inc. [Legal](https://www.chaossearch.io/legal)

Elasticsearch, Logstash, and Kibana are trademarks of Elasticsearch B.V., registered in the U.S. and in other countries. Elasticsearch B.V. and ChaosSearch®, Inc., are not affiliated. Equifax is a registered trademark of Equifax, Inc.

Contact Us

Phone: [(800) 216-0202](tel:+8002160202)

Email: [teamchaos@chaossearch.io](mailto:teamchaos@chaossearch.io)

Follow Us

- <https://twitter.com/CHAOSSEARCH>
- <https://www.linkedin.com/company/chaossearch>
- <https://www.youtube.com/@chaossearch-io>
- <https://datalegendspodcast.com>

```json
{
  "@context" : "https://schema.org",
  "@type" : "VideoObject",
  "description" : "Business intelligence from your VPC Flow Logs - From Log Data to Visionary Clarity",
  "name" : "Business intelligence from your VPC Flow Logs - From Log Data to Visionary Clarity",
  "thumbnailUrl" : "https://i.vimeocdn.com/video/1007692891_100x75.jpg?r=pad",
  "transcript" : "[MUSIC PLAYING] DAVID: Hello, everyone. And welcome to today's webinar, Business Intelligence from your VPC Flow Logs. This is part three of our three-part webinar series entitled From Log Data to Visionary Clarity. Before I introduce today's speakers, I'd like to remind you that we will be taking questions at the end of the session. So please use the Questions panel to submit any questions or comments that you have. Our speakers will get to them at the end of the presentation. And now, without further ado, I'd like to introduce you to today's speakers, Stephen Coveney, senior account executive from ChaosSearch, and Kevin Davis, director of solution architecture. And now, Steve, take it away. STEPHEN COVENEY: Thank you, David. ChaosSearch allows our clients to know better. We have created the ability to provide you insights at scale, allowing you to have unlimited data retention while also driving massive savings. Every company and all businesses want to use their data to know better. But as volumes rise with the velocity and the variety increasing, legacy tools fail and it no longer becomes affordable, leaving teams at a loss of insights. So what if you could search and query all of your data at a massive scale while doing it at a cost savings of 80% compared to what you were currently paying today, all while using the same tools that you're using today and that your teams are used to using on a daily basis. This allows your team to know better-- insights at scale, unlimited data retention using your tools today, all at a significant cost savings. Here is an example of log analytics, which we are using as our first use case. You or some of your teams may be using some form of elasticsearch, whether it is AWS Elasticsearch, Elastic.co, or a hosted solution. It's very common to have to find yourself making some sort of trade-off. And it can be between limiting data retention, being selective over the type of data that you are using, selling off or splitting up use cases with having to scale out your environment. It can be really challenging, expensive, and could cause downtime and instability within your environment. These are the limits everyone sees with this type of stack. Because in the end, it's the Lucene database underneath. What we do at ChaosSearch is simple. Just put your data into your cloud object storage. And in this case, we are showing AWS S3. With our published APIs for you to use with your existing tools, you can go after all of your data. And in just one minute, we will show you how you can do this while working with your VPC Flow Logs. This gives you the unlimited data retention. Because with us, there's no data retention costs. You can keep as much as you want in your S3. Or you can move it off into a glacier. And this is where you can see those cost savings, upwards of 80%. We're able to provide this as a managed service. So that gives you 99.99% of uptime. And you end up with one unified data lake and one repository for your cloud object storage. Now, I have a couple of examples of how we are helping our clients know better today. HubSpot was faced with shortened data retention versus cost versus scaling challenges. ChaosSearch has allowed them to reduce their cost while extending their retention far beyond the five-day window that they were previously working with. So we reduced their costs and extended their retention window. And then we have Armor Security. They provide their customers with a customer-facing log analytics platform that was previously also based on elasticsearch. But now, by using ChaosSearch, their clients have long-term visibility and the ability to be flexible with their data retention plans. Now, here's a bit more on how we are able to do this. It's done in just minutes and in three easy steps-- one, store. two, connect; three, analyze. First step, store your data and cloud object storage. Storage your data in S3. It's your bucket, not ours. Step two, connect to ChaosSearch. ChaosSearch integrates directly with your cloud object storage through a read-only policy. We do not move your data, store, or transform, or modify your data in any way. We leave it right in place and write the indices back to your S3 environment. You can set up for a static index, live index, or real-time. We have built-in schema detection and normalizations. All transforms are done virtually and instantly through our data refinery. Step three, analyze. We have published the elasticsearch API so that you can use the tools that you're used to using on a daily basis. Or you can use a Kibana, which we include in our service package as well. Now I'd like to pass it over to my colleague, Kevin Davis, who is our director of solutions architecture. KEVIN DAVIS: Thanks, Steve. Thank you, everyone, for taking the time today. Before we log into ChaosSearch, we'll make a couple of assumptions like we have in the past couple of webinars. Those assumptions are we have VPC Flow Logs set up and enabled with AWS. And we're sending those logs directly to an S3 bucket in ChaosSearch. When first log in to ChaosSearch, we're brought to the main Storage section. And then from there, we can visualize and see all of our S3 buckets on the left-hand side. The four main sections of ChaosSearch are the Storage, Refinery, Analytics, and Dashboard sections. Storage is allowing you to collect and index all of your logs. Refinery, where you can do transformations and define and set index patterns on those indices. Analytics, where you have the ability to search your data through our Kibana interface. And the Dashboard section, which is a high-level management dashboard for all of our customers to see exactly what activity has been going on, how much data has been indexed, and the user activity within the platform itself. To get started with ChaosSearch, we'll want to go ahead and create an IAM role and policy. That policy will give us the read-write access that we need to appropriately index and also search that data stored in S3. To create an index within ChaosSearch, we want to select the specific S3 bucket. We also have an opportunity to run, and discover, and catalog all of the data within the S3 itself. Here, this high-level snapshot shows us the number of files, the total data size within that S3 bucket, when it was created, and some other information on the file types, last modifications, and any duplicates of those files itself. Creating an object group is very simple. We'll go ahead and click Create Object Group in the top-right corner. And then we can start to group our files in a couple of different ways. The most common approach our customers take is to utilize this Prefix filter. So here, if I input my VPC Flow Logs, I can also look at all of VPC Flow Logs for 2020. Or I can remove that and then gather all of the VPC Flow Logs through that prefix itself. From there, I can input a regex, and do some more insight, and gather logs and group them based on that regex filter for the file name. I can also use this Object filter here to define a tag or any of the metadata on those objects itself. It's a group. And again, start indexing those files. In this particular case, I'll use these VPC Flow Logs in this particular section. And then from there, when we click Next, we'll see the system automatically identifies the schema we're working with, whether or not the files are compressed or uncompressed. And then in the case of common log formats, we'll also input the appropriate regex so we can parse this data properly. Within the platform, we also have immediate validation on whether or not the capture groups that we define are accurate-- and, again, that the data is parsed correctly and as needed. We can change any of these capture groups at any point. Once we click the check bar, we can, again, validate a format or preview, understand the specific columns and the fields within those columns that we'll be working with, and then from there, finalize the step in the object group creation. Here, we'll go ahead and give this object group a name. We can define whether or not we want live indexing. For ChaosSearch, there are three types of approaches. The one I'm going through and creating right now is a static index. It takes all of the files from a point in time from the very last upload to the most recent upload, those files. The other option is a live index through AWS SQS messages where any new event put into an S3 bucket will send a message directly to the SQS queue that ChaosSearch has the rights to read from. We also have the ability to utilize real-time indexing, either through the elasticsearch bulk API or through a logstash configuration. Again, in this example, we'll just go with a static index. I also have the ability to define a retention policy, allowing me to configure how long data should be indexed. And anything outside of the threshold that I define will be removed directly from that ChaosSearch index that's stored in your S3 bucket. Of course, the power of ChaosSearch is not having to set retention policies and then locking all of the data within S3 to make it searchable whenever it's absolutely needed, allowing you to know better. Once I click Create, I'll then be able to go ahead and start the indexing process by clicking Start Index. From here, ChaosSearch will build out the entire mapping within the index structure. It will continuously add new fields as they're added into the schema and will also assign the appropriate type to those fields. As data continues to index, we'll have visibility into how much data has been indexed to that point, the last index time, and the ability to define and change the retention policy should that ever be needed. You'll also have the ability to update and change the retention policy at any point should you ever need to. Now that the index has completed, we can move directly into the Refinery, again allowing us to do transformations and defining an index pattern on the index itself. Once I'm in the refinery, here, I'll have visibility into all the object groups that have been created. From here, I can go and create a view. Views within ChaosSearch allow me to do a couple of different things. I can take independent object groups, merge them together into one single index to search across within Kibana. Or I can select the specific one that we were working with in the Storage section. I can define an index window, allowing me to capture a sliding window of time on what data is available. And then at this point, I can go ahead and see any of the specific fields already indexed that are now available for transformation. In the event that I want to take some of this information-- specifically, the source address-- I can also define whether or not I want to treat those IPs as an IP address or treat it as a geopoint for any additional visualizations within the platform. Again, here, we can load this. And we can do transformations on these fields should we need to. But in this case, we won't go ahead and take that action. We'll simply just define this IP address to be treated as an IP. So when we're searching that data in Kibana, it'll do so. The next step is to define that time period for the index itself. Once I've selected the appropriate time stamp, I can finish selecting and creating this view and then start searching on this data directly in Kibana itself. I also have the ability to define whether or not I want this index to be cacheable, allowing the search performance to improve based on data being cached during search time. I'll go ahead and give this view a new name. We'll call it VPC Flow Logs Webinar. I can also select whether or not I want this index to have caching enabled. Caching is specific to searching only. And it will improve the search latency for any of our end users. I can also define whether or not I want this index to have Case Insensitive Search enabled as well. In this example, we'll leave it undefined. I'll create the view. And from here, I'll have the ability to drill into any of the summary information on that data. Again, we can see the source address is being treated as an IP and not as a string within ChaosSearch. And then, again, we can also update and modify that index window at any point for that sliding scale of time to search against. We also have visibility and understanding on whether or not this index is cacheable and, again, whether or not we define Case Insensitive Searching on that as well. I'll go ahead and make this a favorite. Now, moving into the Analytics section, we'll be brought directly into the management aspect of Kibana. Our version of Kibana is through open distro. And we're on one of the recent releases of 7.4.2. With that, we have the ability to set and create alerts as well as define role-based access control groups and define writes on which users have different access to the different sections of the platform as well as which visualizations, dashboards, and indices are available for them to search against. Moving into the Discover section, we'll see we've been directly dropped into the Webinar VPC Flow Log example that we were working with. And we can now see all of the recent events from our VPC Flow Logs. We'll have the ability to drill into any of the specific rows and see exactly what information we're looking for, any Accept actions, the status of the logs themselves-- more importantly, the source address, source ports, and then any additional information on bytes sent. From here, we can continuously drill into and create different views on this data to understand exactly what's been going on for the connections between our different VPCs. As I go through and continuously add this information, we can also add it into the filter and search on that information across time. And from there, we can use the left-hand menu to drill into additional visualizations and information to see exactly what's been going on. From here, I can go ahead and drill into any of the pieces on the left-hand side again to see what's been going on and what's been taking place. I can quickly drill into this information, create different visualizations, and see what trends might occur over that time period that I'm searching against-- again, keeping that specific address that we're focused on pinned on that search itself. From here, I can go ahead and modify the visualization and add in new terms. Or I can simply modify this to have a date histogram, add that term in again that we were working with. And see, again, trends over time and the number of Accepts for that time period. Here, we can see those trends over time on a number of Accepts within the indice itself. Now, all of these pieces of information for the visualization are great. But now, with the three-part webinar that we've gone through, we can now incorporate all of our ELB logs, our cloud trail logs, and these VPC Flow Logs into one single dashboard specific to the three areas that we've focused in this webinar series-- our cloud trail logs, our ELB logs, and our VPC Flow Logs, giving us full visibility across our AWS infrastructure and understanding which API calls are being made, load balancers that are being hit, and the traffic connectivity within our VPCs itself. Here, in this dashboard, we can see all of that information visible to us. We can have it up and available at any point for all of our users. And again, we can drill into and search on any of the relevant information within the visual itself, continuously iterating and understanding which activity has been going on. Again, the final section within the platform is the Dashboard. This dashboard gives us high-level insight into the total data index, how many indices we have on a daily basis, the total number of queries that have been executed, as well as the average search duration per query. Here, in this section of indexes, we'll break out the individual indices themselves, show you the amount of data per index and the last index time. Here in the Query section, we'll break out the individual queries themselves. We'll show the specific index that was searched on, the user that searched on that index, and the search response time for the search that was executed. Down here, we'll break out the individual events, giving us insight into any potential parse errors with the regex itself. We'll also show the index bytes over time and any additional events over time. To get started with ChaosSearch, you can go to the ChaosSearch website, download a free trial. You can also join us at Reinvent, where we'll have a virtual booth. Once we've gone and signed up for a ChaosSearch account, we can go ahead and start adding our users and testing out the three different workflows that we've gone through in this webinar series. I want to thank you for the time today. And I'll hand it back over to David to take any questions and answers. ",
  "uploadDate" : "2021-03-25T17:21:49.000-04:00"
}
```