297 lines
16 KiB
Markdown
Executable File
297 lines
16 KiB
Markdown
Executable File
---
|
|
page-title: "elasticdump/elasticsearch-dump - Docker Image | Docker Hub"
|
|
url: https://hub.docker.com/r/elasticdump/elasticsearch-dump
|
|
date: "2024-09-26 11:11:29"
|
|
---
|
|
Tools for moving and saving indices.
|
|
|
|
Elasticdump works by sending an `input` to an `output`. Both can be either an elasticsearch URL or a File.
|
|
|
|
If Elasticsearch is not being served from the root directory the `--input-index` and `--output-index` are required. If they are not provided, the additional sub-directories will be parsed for index and type.
|
|
|
|
If you prefer using docker to use elasticdump, you can download this project from docker hub:
|
|
|
|
The file format generated by this tool is line-delimited JSON files. The dump file itself is not valid JSON, but each line is. We do this so that dumpfiles can be streamed and appended without worrying about whole-file parser integrity.
|
|
|
|
```
|
|
elasticdump: Import and export tools for elasticsearch
|
|
version: %%version%%
|
|
|
|
Usage: elasticdump --input SOURCE --output DESTINATION [OPTIONS]
|
|
|
|
--input
|
|
Source location (required)
|
|
--input-index
|
|
Source index and type
|
|
(default: all, example: index/type)
|
|
--output
|
|
Destination location (required)
|
|
--output-index
|
|
Destination index and type
|
|
(default: all, example: index/type)
|
|
--overwrite
|
|
Overwrite output file if it exists
|
|
(default: false)
|
|
--limit
|
|
How many objects to move in batch per operation
|
|
limit is approximate for file streams
|
|
(default: 100)
|
|
--size
|
|
How many objects to retrieve
|
|
(default: -1 -> no limit)
|
|
--concurrency
|
|
The maximum number of requests the can be made concurrently to a specified transport.
|
|
(default: 1)
|
|
--concurrencyInterval
|
|
The length of time in milliseconds in which up to <intervalCap> requests can be made
|
|
before the interval request count resets. Must be finite.
|
|
(default: 5000)
|
|
--intervalCap
|
|
The maximum number of transport requests that can be made within a given <concurrencyInterval>.
|
|
(default: 5)
|
|
--carryoverConcurrencyCount
|
|
If true, any incomplete requests from a <concurrencyInterval> will be carried over to
|
|
the next interval, effectively reducing the number of new requests that can be created
|
|
in that next interval. If false, up to <intervalCap> requests can be created in the
|
|
next interval regardless of the number of incomplete requests from the previous interval.
|
|
(default: true)
|
|
--throttleInterval
|
|
Delay in milliseconds between getting data from an inputTransport and sending it to an
|
|
outputTransport.
|
|
(default: 1)
|
|
--debug
|
|
Display the elasticsearch commands being used
|
|
(default: false)
|
|
--quiet
|
|
Suppress all messages except for errors
|
|
(default: false)
|
|
--type
|
|
What are we exporting?
|
|
(default: data, options: [settings, analyzer, data, mapping, policy, alias, template, component_template, index_template])
|
|
--filterSystemTemplates
|
|
Whether to remove metrics-*-* and logs-*-* system templates
|
|
(default: true])
|
|
--templateRegex
|
|
Regex used to filter templates before passing to the output transport
|
|
(default: ((metrics|logs|\\..+)(-.+)?)
|
|
--delete
|
|
Delete documents one-by-one from the input as they are
|
|
moved. Will not delete the source index
|
|
(default: false)
|
|
--searchBody
|
|
Preform a partial extract based on search results
|
|
when ES is the input, default values are
|
|
if ES > 5
|
|
`'{"query": { "match_all": {} }, "stored_fields": ["*"], "_source": true }'`
|
|
else
|
|
`'{"query": { "match_all": {} }, "fields": ["*"], "_source": true }'`
|
|
[As of 6.68.0] If the searchBody is preceded by a @ symbol, elasticdump will perform a file lookup
|
|
in the location specified. NB: File must contain valid JSON
|
|
--searchWithTemplate
|
|
Enable to use Search Template when using --searchBody
|
|
If using Search Template then searchBody has to consist of "id" field and "params" objects
|
|
If "size" field is defined within Search Template, it will be overridden by --size parameter
|
|
See https://www.elastic.co/guide/en/elasticsearch/reference/current/search-template.html for
|
|
further information
|
|
(default: false)
|
|
--searchBodyTemplate
|
|
A method/function which can be called to the searchBody
|
|
doc.searchBody = { query: { match_all: {} }, stored_fields: [], _source: true };
|
|
May be used multiple times.
|
|
Additionally, searchBodyTemplate may be performed by a module. See [searchBody Template](#search-template) below.
|
|
--headers
|
|
Add custom headers to Elastisearch requests (helpful when
|
|
your Elasticsearch instance sits behind a proxy)
|
|
(default: '{"User-Agent": "elasticdump"}')
|
|
Type/direction based headers are supported .i.e. input-headers/output-headers
|
|
(these will only be added based on the current flow type input/output)
|
|
--params
|
|
Add custom parameters to Elastisearch requests uri. Helpful when you for example
|
|
want to use elasticsearch preference
|
|
|
|
--input-params is a specific params extension that can be used when fetching data with the scroll api
|
|
--output-params is a specific params extension that can be used when indexing data with the bulk index api
|
|
NB : These were added to avoid param pollution problems which occur when an input param is used in an output source
|
|
(default: null)
|
|
--sourceOnly
|
|
Output only the json contained within the document _source
|
|
Normal: {"_index":"","_type":"","_id":"", "_source":{SOURCE}}
|
|
sourceOnly: {SOURCE}
|
|
(default: false)
|
|
--ignore-errors
|
|
Will continue the read/write loop on write error
|
|
(default: false)
|
|
--scrollId
|
|
The last scroll Id returned from elasticsearch.
|
|
This will allow dumps to be resumed used the last scroll Id &
|
|
`scrollTime` has not expired.
|
|
--scrollTime
|
|
Time the nodes will hold the requested search in order.
|
|
(default: 10m)
|
|
|
|
--scroll-with-post
|
|
Use a HTTP POST method to perform scrolling instead of the default GET
|
|
(default: false)
|
|
|
|
--maxSockets
|
|
How many simultaneous HTTP requests can we process make?
|
|
(default:
|
|
5 [node <= v0.10.x] /
|
|
Infinity [node >= v0.11.x] )
|
|
--timeout
|
|
Integer containing the number of milliseconds to wait for
|
|
a request to respond before aborting the request. Passed
|
|
directly to the request library. Mostly used when you don't
|
|
care too much if you lose some data when importing
|
|
but rather have speed.
|
|
--offset
|
|
Integer containing the number of rows you wish to skip
|
|
ahead from the input transport. When importing a large
|
|
index, things can go wrong, be it connectivity, crashes,
|
|
someone forgets to `screen`, etc. This allows you
|
|
to start the dump again from the last known line written
|
|
(as logged by the `offset` in the output). Please be
|
|
advised that since no sorting is specified when the
|
|
dump is initially created, there's no real way to
|
|
guarantee that the skipped rows have already been
|
|
written/parsed. This is more of an option for when
|
|
you want to get most data as possible in the index
|
|
without concern for losing some rows in the process,
|
|
similar to the `timeout` option.
|
|
(default: 0)
|
|
--noRefresh
|
|
Disable input index refresh.
|
|
Positive:
|
|
1. Much increase index speed
|
|
2. Much less hardware requirements
|
|
Negative:
|
|
1. Recently added data may not be indexed
|
|
Recommended using with big data indexing,
|
|
where speed and system health in a higher priority
|
|
than recently added data.
|
|
--inputTransport
|
|
Provide a custom js file to use as the input transport
|
|
--outputTransport
|
|
Provide a custom js file to use as the output transport
|
|
--toLog
|
|
When using a custom outputTransport, should log lines
|
|
be appended to the output stream?
|
|
(default: true, except for `$`)
|
|
--transform
|
|
A method/function which can be called to modify documents
|
|
before writing to a destination. A global variable 'doc'
|
|
is available.
|
|
Example script for computing a new field 'f2' as doubled
|
|
value of field 'f1':
|
|
doc._source["f2"] = doc._source.f1 * 2;
|
|
May be used multiple times.
|
|
Additionally, transform may be performed by a module. See [Module Transform](#module-transform) below.
|
|
--awsChain
|
|
Use [standard](https://aws.amazon.com/blogs/security/a-new-and-standardized-way-to-manage-credentials-in-the-aws-sdks/) location and ordering for resolving credentials including environment variables, config files, EC2 and ECS metadata locations
|
|
_Recommended option for use with AWS_
|
|
Use [standard](https://aws.amazon.com/blogs/security/a-new-and-standardized-way-to-manage-credentials-in-the-aws-sdks/)
|
|
location and ordering for resolving credentials including environment variables,
|
|
config files, EC2 and ECS metadata locations _Recommended option for use with AWS_
|
|
--awsAccessKeyId
|
|
--awsSecretAccessKey
|
|
When using Amazon Elasticsearch Service protected by
|
|
AWS Identity and Access Management (IAM), provide
|
|
your Access Key ID and Secret Access Key.
|
|
--sessionToken can also be optionally provided if using temporary credentials
|
|
--awsIniFileProfile
|
|
Alternative to --awsAccessKeyId and --awsSecretAccessKey,
|
|
loads credentials from a specified profile in aws ini file.
|
|
For greater flexibility, consider using --awsChain
|
|
and setting AWS_PROFILE and AWS_CONFIG_FILE
|
|
environment variables to override defaults if needed
|
|
--awsIniFileName
|
|
Override the default aws ini file name when using --awsIniFileProfile
|
|
Filename is relative to ~/.aws/
|
|
(default: config)
|
|
--awsService
|
|
Sets the AWS service that the signature will be generated for
|
|
(default: calculated from hostname or host)
|
|
--awsRegion
|
|
Sets the AWS region that the signature will be generated for
|
|
(default: calculated from hostname or host)
|
|
--awsUrlRegex
|
|
Overrides the default regular expression that is used to validate AWS urls that should be signed
|
|
(default: ^https?:\/\/.*\.amazonaws\.com.*$)
|
|
--support-big-int
|
|
Support big integer numbers
|
|
--big-int-fields
|
|
Sepcifies a comma-seperated list of fields that should be checked for big-int support
|
|
(default '')
|
|
--retryAttempts
|
|
Integer indicating the number of times a request should be automatically re-attempted before failing
|
|
when a connection fails with one of the following errors `ECONNRESET`, `ENOTFOUND`, `ESOCKETTIMEDOUT`,
|
|
ETIMEDOUT`, `ECONNREFUSED`, `EHOSTUNREACH`, `EPIPE`, `EAI_AGAIN`
|
|
(default: 0)
|
|
|
|
--retryDelay
|
|
Integer indicating the back-off/break period between retry attempts (milliseconds)
|
|
(default : 5000)
|
|
--parseExtraFields
|
|
Comma-separated list of meta-fields to be parsed
|
|
--maxRows
|
|
supports file splitting. Files are split by the number of rows specified
|
|
--fileSize
|
|
supports file splitting. This value must be a string supported by the **bytes** module.
|
|
The following abbreviations must be used to signify size in terms of units
|
|
b for bytes
|
|
kb for kilobytes
|
|
mb for megabytes
|
|
gb for gigabytes
|
|
tb for terabytes
|
|
|
|
e.g. 10mb / 1gb / 1tb
|
|
Partitioning helps to alleviate overflow/out of memory exceptions by efficiently segmenting files
|
|
into smaller chunks that then be merged if needs be.
|
|
--fsCompress
|
|
gzip data before sending output to file.
|
|
On import the command is used to inflate a gzipped file
|
|
--s3AccessKeyId
|
|
AWS access key ID
|
|
--s3SecretAccessKey
|
|
AWS secret access key
|
|
--s3Region
|
|
AWS region
|
|
--s3Endpoint
|
|
AWS endpoint can be used for AWS compatible backends such as
|
|
OpenStack Swift and OpenStack Ceph
|
|
--s3SSLEnabled
|
|
Use SSL to connect to AWS [default true]
|
|
|
|
--s3ForcePathStyle Force path style URLs for S3 objects [default false]
|
|
|
|
--s3Compress
|
|
gzip data before sending to s3
|
|
--s3ServerSideEncryption
|
|
Enables encrypted uploads
|
|
--s3SSEKMSKeyId
|
|
KMS Id to be used with aws:kms uploads
|
|
--s3ACL
|
|
S3 ACL: private | public-read | public-read-write | authenticated-read | aws-exec-read |
|
|
bucket-owner-read | bucket-owner-full-control [default private]
|
|
--s3StorageClass
|
|
Set the Storage Class used for s3
|
|
(default: STANDARD)
|
|
--s3Options
|
|
Set all s3 parameters shown here https://docs.aws.amazon.com/AWSJavaScriptSDK/latest/AWS/S3.html#createMultipartUpload-property
|
|
A escaped JSON string or file can be supplied. File location must be prefixed with the @ symbol
|
|
(default: null)
|
|
--s3Configs
|
|
Set all s3 constructor configurations
|
|
A escaped JSON string or file can be supplied. File location must be prefixed with the @ symbol
|
|
(default: null)
|
|
--retryDelayBase
|
|
The base number of milliseconds to use in the exponential backoff for operation retries. (s3)
|
|
--customBackoff
|
|
Activate custom customBackoff function. (s3)
|
|
--tlsAuth
|
|
Enable TLS X509 client authentication
|
|
--cert, --input-cert, --output-cert
|
|
Client certificate file. Use --cert if source and destination are identical.
|
|
Otherwise, use the one prefixed with --input or --output as needed.
|
|
--key, --input-key, --o
|
|
``` |