---
title: "ChaosSearch Data Refinery: transform without reindexing"
description: "Announcing CHAOSSEARCH Data Refinery: Transform Data and Schema on the Fly Without Reindexing"
image: https://www.chaossearch.io/hubfs/chaossearch-blog-banner.svg
---

Revinate leaves their ELK stack behind to find huge gains with ChaosSearch -- Read More!

[Revinate leaves their ELK stack behind to find huge gains with ChaosSearch -- Read More!

](https://www.chaossearch.io/resources/customer-stories/revinate)

[![ChaosSearch](https://www.chaossearch.io/hubfs/2021%20Website/logo.svg) ](https://www.chaossearch.io/)

![Gartner Cool Vendor 2023](https://www.chaossearch.io/hubfs/C2020/Logos/Gartner%20Cool%20Vendor%202023.png)

[![Start Free Trial](https://no-cache.hubspot.com/cta/default/4020721/a1eeee12-14a2-41d7-a38b-11e32980b916.png)](https://cta-redirect.hubspot.com/cta/redirect/4020721/a1eeee12-14a2-41d7-a38b-11e32980b916)

## Interested in scheduling a brief intro call to see how ChaosSearch can accelerate your analytics?

Yes - show me the calendar!

## [ChaosSearch Blog](https://www.chaossearch.io/blog)

6 MIN READ

# ChaosSearch Data Refinery: transform without reindexing

By [Pete Cheslock](https://www.chaossearch.io/blog/chaossearch-data-refinery-transform-without-reindexing#authorBlock) on Aug 6, 2019

ChaosSearch Data Refinery: transform without reindexing

4:57

Traditional databases suffer a problem when ingesting data. They operate on a schema-on-write approach where data indexed must have a predefined schema as you ingest your data into the database. This schema-on-write model means that you need to take time in advance to dive into your data and understand what is there, and then process your data in advance to fit the defined schema. This data cleaning and data processing can be time-consuming and costly, not only with the engineering time to build but also the computing costs associated with it.

Now for many companies’ log data, they can adjust code within their software to clean the data to prepare it for indexing, but what about data coming from sources that you can't control? Various Amazon cloud services can generate vast amounts of log and event data and send these logs to your Amazon S3 buckets, but you can't modify the data before it lands in S3.[ You'd have to build various Lambda functions or other post-processing tasks](https://www.chaossearch.io/blog/reduce-complexity-and-quickly-search-amazon-cloudfront-logs-in-amazon-s3) to parse and convert the data into an appropriate format after it's been written to your bucket.

The ultimate goal of **CHAOS**SEARCH is to help our customers get quick insights into their data without ever having to move their data out of Amazon S3. That's why I'm incredibly excited to announce the **CHAOS**SEARCH Data Refinery today. Now you can transform your already indexed data, creating new fields that can be searched for as well as adjusting the objects' schema on the fly. No longer do you need to extract data and transform it (ETL) just to get better insights into your logs, you can now do these virtual transformations as different data views.

> With **CHAOS**SEARCH, we are able to quickly create materialized views that effectively and accurately parse our various log formats.
> 
> Transeo

Legacy search technologies like Elasticsearch operate on a schema-on-write approach, and to change the schema for any data field, you would need to reindex that data. Add a new field or change a field type? Change your mapping, reindex your data. Find a mistake in your mapping, change your mapping, reindex your data. Maybe not a big deal if you only have a few hundred GB of data total, but cost and time prohibitive when dealing with 10s or 100s of TBs of data.

With this new release of **CHAOS**SEARCH not only can you transform the field into multiple new fields to query and aggregate, but you can also adjust the schema, changing fields from strings to integers, allowing you to search on ranges. An entirely new view into your data, created within seconds, all without ever having to spend time and money reindexing your data.

One example we see from customers is with Amazon ELB logs. These logs have a unique format.

`timestamp elb client:port backend:port request_processing_time backend_processing_time response_processing_time elb_status_code backend_status_code received_bytes sent_bytes "request" "user_agent" ssl_cipher ssl_protocol`

![Announcing CHAOSSEARCH Data Refinery: Transform Data and Schema on the Fly Without Reindexing](https://www.chaossearch.io/hubfs/Imported_Blog_Media/data-refinery-1-1-1024x332.png)

The annoying part of these logs is that both the client and the instance IP addresses include the port within the string. If you had already indexed this data or didn't want to spend any time in advance to parse it, you'd be stuck with the inability to analyze this data quickly.

In this example, you can see the initial indexing process identified both the client and backend IP as a string — **CHAOS**SEARCH automatically infers schema on your data to speed up your time to insights.

![Announcing CHAOSSEARCH Data Refinery: Transform Data and Schema on the Fly Without Reindexing](https://www.chaossearch.io/hubfs/Imported_Blog_Media/data-refinery-2-1024x384.png)

When I go into the **CHAOS**SEARCH embedded Kibana interface and search for an IP address...

`client_ip : "235.237.171.163"`

You can see no results:

![Announcing CHAOSSEARCH Data Refinery: Transform Data and Schema on the Fly Without Reindexing](https://www.chaossearch.io/hubfs/Imported_Blog_Media/data-refinery-3-1024x246.png)

We'd have to know the port as well to find it, OR we can add a wildcard to the end of the query to find it.

![Announcing CHAOSSEARCH Data Refinery: Transform Data and Schema on the Fly Without Reindexing](https://www.chaossearch.io/hubfs/Imported_Blog_Media/data-refinery-4-1024x290.png)

However, this isn't a great user experience — and this also makes it nearly impossible to see the most and least frequent IP addresses because the port number exists in all aggregations.

![Announcing CHAOSSEARCH Data Refinery: Transform Data and Schema on the Fly Without Reindexing](https://www.chaossearch.io/hubfs/Imported_Blog_Media/data-refinery-5-1024x392.png)

With the new **CHAOS**SEARCH Data Refinery, we can transform this field in seconds using a regular expression, and split this into a string field and a number field.

![Announcing CHAOSSEARCH Data Refinery: Transform Data and Schema on the Fly Without Reindexing](https://www.chaossearch.io/hubfs/Imported_Blog_Media/data-refinery-6-1024x481.png)

Now we have 2 new fields:

`client_ip_address client_port`

Back in Kibana, I can immediately see a new view created that includes the new fields with the updated schema.

![Announcing CHAOSSEARCH Data Refinery: Transform Data and Schema on the Fly Without Reindexing](https://www.chaossearch.io/hubfs/Imported_Blog_Media/data-refinery-7.png)

Now I can run a search JUST for the IP address I'm looking for — no wildcards required.

![Announcing CHAOSSEARCH Data Refinery: Transform Data and Schema on the Fly Without Reindexing](https://www.chaossearch.io/hubfs/Imported_Blog_Media/data-refinery-8-1024x507.png)

Also, most importantly I can run aggregations on the new data fields that I've created and get quick answers to my questions, all without having to write code to reformat my data, and all without having to spend time and money reindexing my data. I can use these virtual transformations to potentially mask out PII or other identifiable customer data before I run various reports. 

**CHAOS**SEARCH is helping customers get immediate insights into their data without having to do anything in advance. Dump your data into Amazon S3, index it one time with **CHAOS**SEARCH, and transform and modify it in endless ways. Store everything. Ask anything.

[Reach out for a trial](https://www.chaossearch.io/trial/) and try this new feature today!

   [News](https://www.chaossearch.io/blog/tag/news)

### About the Author, Pete Cheslock

![Pete Cheslock](https://www.chaossearch.io/hubfs/ChaosSearch_September2019/Images/pete-150x150.png)

FOLLOW ME ON:

[Pete Cheslock's LinkedIn ](https://www.linkedin.com/in/petecheslock/)

 Pete Cheslock was the VP of Product for ChaosSearch, where he was brought on as one of the founding executives. In his role, Pete helped to define the go-to-market strategy and refine product direction for the initial ChaosSearch launch. To see what Pete’s up to now, connect with him on LinkedIn. [More posts by Pete Cheslock](https://www.chaossearch.io/blog/author/pete-cheslock)

## You may also like

## Future-Proof Your Analytics at Scale

[![Get a Demo](https://no-cache.hubspot.com/cta/default/4020721/0ae4b1b0-9ec0-4169-80cc-64fda5fd56db.png)](https://cta-redirect.hubspot.com/cta/redirect/4020721/0ae4b1b0-9ec0-4169-80cc-64fda5fd56db)

©2024, ChaosSearch®, Inc. [Legal](https://www.chaossearch.io/legal)

Elasticsearch, Logstash, and Kibana are trademarks of Elasticsearch B.V., registered in the U.S. and in other countries. Elasticsearch B.V. and ChaosSearch®, Inc., are not affiliated. Equifax is a registered trademark of Equifax, Inc.

Contact Us

Phone: [(800) 216-0202](tel:+8002160202)

Email: [teamchaos@chaossearch.io](mailto:teamchaos@chaossearch.io)

Follow Us

- <https://twitter.com/CHAOSSEARCH>
- <https://www.linkedin.com/company/chaossearch>
- <https://www.youtube.com/@chaossearch-io>
- <https://datalegendspodcast.com>

```json
{
  "@context" : "https://schema.org",
  "@type" : "BlogPosting",
  "author" : {
    "@type" : "Person",
    "name" : "Pete Cheslock",
    "url" : "https://www.chaossearch.io/blog/author/pete-cheslock"
  },
  "dateModified" : "2024-10-31T22:48:32.947Z",
  "datePublished" : "2019-08-06T17:48:16.000Z",
  "headline" : "ChaosSearch Data Refinery: transform without reindexing",
  "image" : [ "https://www.chaossearch.io/hubfs/chaossearch-blog-banner.svg" ],
  "mainEntityOfPage" : {
    "@id" : "https://www.chaossearch.io/blog/chaossearch-data-refinery-transform-without-reindexing",
    "@type" : "WebPage"
  },
  "publisher" : {
    "@type" : "Organization",
    "logo" : {
      "@type" : "ImageObject",
      "url" : "https://www.chaossearch.io/hubfs/2021%20Website/logo.svg"
    },
    "name" : "CHAOSSEARCH"
  }
}
```