# Monitoring::Sneck

## SYNOPSIS

```
sneck -u [-C <cache file>] [-f <config file>] [-p] [-i] [-d] [-q] [-l] [-r]

sneck -c [-C <cache file>] [-b]

sneck [-f <config file>] [-p] [-i] [-r]

sneck -t [-f <config file>]
```

## FLAGS

### -f config_file

The config file to use.

Files ending in .yaml or .yml are read as YAML, which needs YAML::XS.

Default :: /usr/local/etc/sneck.conf

### -p

Pretty it in a nicely formatted format.

### -C cache_file

The cache file to use.

Default :: /var/cache/sneck.cache

A secondary cache file based on this name is also created. By default
it is /var/cache/sneck.cache.snmp and always holds the GZip+BASE64
compressed version.

### -u

Update the cache file. Will also print the was written to it.

### -c

Print the cache file. Please note that -p or -i won't affect
this as this flag only reads/prints the cache file.

### -b

When used with -c, print the LibreNMS style GZip+BASE64 compressed
cache instead.

### -i

Includes the config file used.

### -d

Print debugging info if called with -u.

### -q

Don't print the results for -u. Exit quietly.

### -l

Enable locking for -u so more than one instance can't run at a time.

### -r

Run any restarts whose checks failed. Without this, restarts are only
reported as `restarts disabled`. Meant to be used with -u from cron.
Never use it from snmpd.

Restart state, used for `min_interval` and `max_retries`, is kept in the
cache file name with `.restarts` added, so by default
/var/cache/sneck.cache.restarts.

Use -l with this to make sure two runs can't restart things at the same
time.

### -t

Test the config file. Prints any errors and warnings, then exits 0 if
there are no errors and 1 if there are or the file can't be read.
Warnings, such as a check using a undefined variable, do not affect the
exit code.

## CONFIG FORMAT

The sneck format is described below. Files ending in .yaml or .yml
use the YAML format instead, described under YAML CONFIG.

Each line has leading spaces and tabs removed before it is looked at. A
trailing \r is also removed, so files with CRLF line endings work.

Blank lines are ignored.

Lines starting with # are comments and are ignored.

- `env NAME=value` :: Sets a environment variable. The `env` is case
  insensitive. The value may be empty. These may be set more than once,
  with the last one winning. They are only applied if the whole config
  is valid.

- `NAME=value` :: A variable. The name is before the first =, the value
  is everything after it. The value may be empty.

- `name|command` :: A check. The command is everything after the first
  | with leading whitespace removed. It may not be empty.

- `%name|command` :: A debug check. Same as a check, but not counted
  towards any of the counts. It exists purely for debugging. The leading
  % is not part of the name, so a check and a debug check may share a
  name.

- `@name|options|command` :: A restart. Options are space separated
  `key=value`, with `checks` and `depends` being comma separated lists.
  Values containing spaces may be quoted with `"` or `'`. See RESTARTS.

Names are made up of `A-Z`, `a-z`, `0-9`, and `_`.

Any other sort of line is an error. Every bad line is reported, along
with its line number.

Variables are used in commands in the form `%NAME%`. A reference to a
variable that is not defined is left as written and produces a warning,
as it may just be part of the command, such as `date +%Y%m%d`. Use `-t`
to see warnings.

Variable, check, debug check, and restart names may not be redefined.

## EXAMPLE CONFIG

```
env PATH=/sbin:/bin:/usr/sbin:/usr/bin:/usr/local/sbin:/usr/local/bin
# this is a comment
GEOM_DEV=foo
geom_foo|/usr/local/libexec/nagios/check_geom mirror %GEOM_DEV%
does_not_exist|/bin/this_will_error yup... that it will
    
does_not_exist_2|/usr/bin/env /bin/this_will_also_error

#includes route info
%routes|netstat -rn
```

The first line sets the %ENV variable PATH.

The second is ignored as it is a comment.

The third sets the variable GEOM_DEV to 'foo'

The fourth creates a check named geom_foo that calls check_geom_mirror
with the variable supplied to it being the value specified by the
variable GEOM_DEV.

The fith is a example of an error that will show what will happen when
you call to a file that does not exit.

The sixth line will be ignored as it is blank.

The seventh is a example of another command erroring.

When you run it, you will notice that errors for lines 4 and 5 are
printed to STDERR. For this reason you should use '2> /dev/null' when
calling it from snmpd or '2> /dev/null > /dev/null' when calling from
cron. 

## YAML CONFIG

YAML configs need YAML::XS, which is optional and only loaded when a
YAML config is used.

```
env:
  PATH: /sbin:/bin:/usr/sbin:/usr/bin:/usr/local/sbin:/usr/local/bin
vars:
  GEOM_DEV: foo
checks:
  geom_foo: /usr/local/libexec/nagios/check_geom mirror %GEOM_DEV%
  does_not_exist: /bin/this_will_error yup... that it will
debugs:
  routes: netstat -rn
```

- `env` :: Environment variables to set, in sorted name order. Only
  applied if the whole config is valid.

- `vars` :: Variables.

- `checks` :: Checks, with the command as the value.

- `debugs` :: Debug checks, with the command as the value.

- `restarts` :: Restarts, each a mapping of options plus `command`.
  See RESTARTS.

Any other top level key is an error. Any section may be left out or
empty.

Names, variables, and undefined variable warnings work the same as the
sneck format. Errors and warnings point to a path, such as
`checks.geom_foo`, as YAML::XS does not give line numbers.

Values must be strings or numbers. An empty value is an empty string
for `env` and `vars` and an error for `checks` and `debugs`.

Some things to watch for.

- Unquoted `true` becomes 1 and `false` becomes an empty string. Quote
  values like those.

- A value starting with `%` must be quoted, as `%` can't start a plain
  YAML value.

- A value containing `: ` or ending in `:`, such as `-c 1:`, must be
  quoted, or it is read as a mapping.

- Duplicate keys are not caught. YAML::XS silently keeps the last one.

## RESTARTS

A restart is a command run when enough of the checks it watches fail.
They only run with -r. Otherwise they are just reported.

```
http_check|/usr/local/libexec/nagios/check_http -H localhost
php_check|/usr/local/libexec/nagios/check_procs -C php-fpm -c 1:
@php_fpm|checks=php_check|/usr/sbin/service php_fpm restart
@httpd|checks=http_check,php_check threshold=2 depends=php_fpm cascade=1 timeout=60|/usr/sbin/service apache24 restart
```

```
restarts:
  php_fpm:
    command: /usr/sbin/service php_fpm restart
    checks: [php_check]
  httpd:
    command: /usr/sbin/service apache24 restart
    checks: [http_check, php_check]
    threshold: 2
    depends: [php_fpm]
    cascade: true
    timeout: 60
```

- `checks` :: Required. The checks to watch. Debug checks can't be
  watched.

- `threshold` :: How many of the watched checks must fail for it to
  trigger. Default :: 1

- `depends` :: Other restarts this one depends on. When both trigger,
  the ones depended on run first. A depend that didn't trigger is
  assumed fine and isn't run. If a depend runs and fails, this one is
  skipped. Cycles are errors.

- `cascade` :: If 1, this also runs when any of its depends ran without
  failing, even if its own threshold wasn't met. Default :: 0

- `ignore_unknown` :: If 0, unknown counts as failed. Default :: 1

- `ignore_errored` :: If 0, errored counts as failed. This is any exit
  other than 0 to 3, dying on a signal, or not being able to run the
  check. Default :: 1

- `min_interval` :: Minimum seconds between runs. 0 turns it off.
  Default :: 180

- `max_retries` :: How many times in a row it will run for its threshold
  before giving up until its checks recover. 0 means always retry. Runs
  from cascade don't count. Default :: 0

- `timeout` :: Seconds to wait on the command before giving up on it.
  On timeout its output pipes are closed and it is left running. Nothing
  is sent to it, but if it writes again it gets SIGPIPE, or whatever it
  does when the reader goes away. A timeout counts as failed, with an
  exit of -1. Default :: 30

- `not_every` :: A maintenance window, as a five field cron spec. While
  local time matches it, the restart doesn't run. Numbers only, no names
  like `sat` or `jan`. Sunday is 0 or 7. If both day of month and day of
  week are given, either matching is enough, same as cron. Restarts
  depending on one held back by this act as if it didn't trigger. Must
  be quoted, such as `not_every="* 2-3 * * 0"`, or in YAML
  `not_every: '* 2-3 * * 0'`. Default :: none

Critical always counts as failed. Ok and warning never do.

A restart that ran and failed sets `.data.alert` and adds a line to
`.data.alertString`, along with any output from the command.

## USAGE

snmpd should be configured as below.

```
extend sneck /usr/bin/env PATH=/sbin:/bin:/usr/sbin:/usr/bin:/usr/local/sbin:/usr/local/bin /usr/local/bin/sneck -c
```

Then just setup a entry in like cron such as below.

```
*/5 * * * * /usr/bin/env PATH=/sbin:/bin:/usr/sbin:/usr/bin:/usr/local/sbin:/usr/local/bin  /usr/local/bin/sneck -u 2> /dev/null > /dev/null
```

Most likely want to run it once per polling interval.

You can use it in a non-cached manner with out cron, but this will result in a
longer polling time for LibreNMS or the like when it queries it.

## RETURN HASH/JSON

The generated JSON/hash is as below in jpath notation.

- .data.alert :: 0/1 boolean for if there is a aloert or not.

- .data.ok :: Count of the number of ok checks.

- .data.warning :: Count of the number of warning checks.

- .data.critical :: Count of the number of critical checks.

- .data.unknown :: Count of the number of unkown checks.

- .data.errored :: Count of the number of errored checks.

- .data.alertString :: The cumulative outputs of anything that
  returned a warning, critical, or unknown.

- .data.vars :: A hash with the variables to use.

- .data.time :: Time since epoch.

- .data.time :: The hostname the check was ran on.

- .data.config :: The raw config file if told to include it.

- $hash{data}{run_time} :: How long it took to run all checks.

- .data.restarted :: Count of the number of restarts ran.

- .data.restart_state_error :: Only present if the restart state file
  could not be read or written.

For the following `$name` is the name of the check ran.

- .data.checks.$name :: A hash with info on the checks ran.

- .data.checks.$name.check :: The command pre-variable substitution.

- .data.checks.$name.ran :: The command ran.

- .data.checks.$name.output :: The output of the check.

- .data.checks.$name.exit :: The exit code.

- .data.checks.$name.error :: Only present it died on a signal or
  could not be executed. Provides a brief description.

- $hash{data}{checks}{$name}{run_time} :: How long it took to run the checks.

For the following `$name` is the name of the debug check ran.

- .data.debugs.$name :: A hash with info on the checks ran.

- .data.debugs.$name.check :: The command pre-variable substitution.

- .data.debugs.$name.ran :: The command ran.

- .data.debugs.$name.output :: The output of the check.

- .data.debugs.$name.exit :: The exit code.

- .data.debugs.$name.error :: Only present it died on a signal or
  could not be executed. Provides a brief description.

 - $hash{data}{checks}{$name}{run_time} :: How long it took to run the debug.

For the following `$name` is the name of the restart. Every restart has
an entry, even when restarts are disabled.

- .data.restarts.$name.triggered :: 0/1 for if its threshold was met.

- .data.restarts.$name.ran :: 0/1 for if it ran.

- .data.restarts.$name.reason :: Why it did or didn't run. One of
  `threshold`, `cascade from $depend`, `cooldown, $N seconds left`,
  `max retries reached`, `skipped, dependency $depend failed`,
  `maintenance window`, `not triggered`, or `restarts disabled`.

- .data.restarts.$name.failed_checks :: The watched checks that failed.

- .data.restarts.$name.threshold :: The threshold.

- .data.restarts.$name.attempts :: Runs for its threshold since its
  checks last recovered.

- .data.restarts.$name.command :: The command pre-variable substitution.

- .data.restarts.$name.ran_command :: The command ran. Only if it ran.

- .data.restarts.$name.output :: The output. Only if it ran.

- .data.restarts.$name.exit :: The exit code, or -1 if it timed out. Only if
  it ran.

- .data.restarts.$name.error :: Only present if it timed out, died on a
  signal, or could not be executed.

- .data.restarts.$name.run_time :: How long it took. Only if it ran.

## INSTALLING

### FreeBSD

```
pkg install p5-JSON p5-JSON-XS p5-File-Slurp p5-Proc-PID-File p5-DateTime-Event-Cron p5-App-cpanminus
cpanminus Monitoring::Sneck
```

### Debian

```
apt-get install libjson-perl libjson-xs-perl libfile-slurp-perl libproc-pid-file-perl libdatetime-event-cron-perl cpanminus
cpanminus Monitoring::Sneck
```

### YAML Support

For YAML configs, also install YAML::XS.

```
# FreeBSD
pkg install p5-YAML-LibYAML

# Debian
apt-get install libyaml-libyaml-perl
```

### From Src

```
perl Makefile.PL
make
make test
make install
```
