Features Overview

Prev Next

Features Overview

A Feature is a number calculated in real time from past events. How many times this card was used in the last hour. How many distinct accounts this device has touched. How long since this email address was first seen, and how far the current login is from the last one.

They exist because the current event is rarely enough on its own. A login from a new device is unremarkable until you know it is the fourth new device on this account today. Features put that history in front of the decision.

Every Feature answers the same question in two parts: which past events count, and what to calculate over them. The criteria settings on this page select the events. The Feature Type decides the calculation.

Features can be calculated on any step of a journey. They are read by Rules, Models and Decision Strategies, and shown in the Sidebar when viewing an event.

Image
Features rendered in the Event Sidebar

Feature Editor

Features are files with the extension ".feature.yaml". The Feature Editor is launched by double clicking on a Feature file within Workflows section of the Darwinium portal.
If no Feature file exists, simply create a new file and give it the extension ".feature.yaml".

image.png

image.png

As the file extension suggests, Features are written in YAML to make them understandable when viewing diffs across different versions.

Notes and Tips
  • Advanced users may choose to edit Features directly in YAML instead of using the Feature Editor.
  • Features can be shared by different Models and Rules (via dependencies).
  • Features are best kept in "Libraries" of common subject matters to help promote re-use. For example, a "Device Features" file could contain features based on Device Signatures and Identifiers, whilst an "Account Features" file could be used to organize Account or User targeted features.

Feature Name

Features have a name to identify them. They are required to be unique per step; name collisions will be called out at build time.

Features output attribute

The result of User defined features are output to key/value pair map attribute:

outcome['CHAMPION'].features.general

The Feature Name is used as a key to retrieve its Feature value:

outcome['CHAMPION'].features.general['account_logins']

The helper function feature is a shortcut to do same thing

feature('account_logins')

Using Features In Rules and Investigations

Features are stored in a name/value pair attribute (also known as dictionary or associative array data type). They can be referenced by name and compared with a value to form a filter (resulting in a Boolean expression).

The following is an example of searching for events where there had been 2 or more logins

feature('account_logins')>=2

Equivalent to:

outcome['CHAMPION'].features.general['account_logins'] >= 2
Investigations Search conditions and Rule conditions are same thing

Anything that is valid search in Investigations can be used directly in a Rule

Feature Types

Counting Features

  • Velocity Of - A count of events with the same attribute value over time.
  • Distinct Count Of - A count of events with unique attribute values over time.
  • Specificity Of - A measurement of how unique a particular attribute value is compared to the full population.
  • Approximate Distinct Count Of - An approx. count of events with unique attribute values over time - not restricted to identifier/subjects.
  • Approximate Total Count Of- An approx. count of events where attribute value is populated (not null) - not restricted to identifier/subjects.
  • Approximate Velocity Of - An approx. count of events matching the criteria, optionally as a fraction of the total rather than a count.

Statistic Features

  • Statistic Of Numeric - The value of a numeric attribute across multiple events into common statistics .(sum/min/max/mean/stddev).
  • Statistic Of Expression - A numeric query language into common statistics (min/max/mean/stddev).
  • Quantile Of Numeric - The probability distribution for a numeric typed attribute.

Time Features

  • Time Between - Elapsed time between multiple events into common statistics (min/max/mean/stddev).
  • Time Since First - Time elapsed from oldest event to now.
  • Time Since Last - Time elapsed from most recent event to now.
  • Probability Of Time - How well activity over a short bucket of time fits the pattern established over a longer periodicity.

Distance Features

  • Distance From - Distance between current and multiple past events into common statistics (min/max/mean/stddev).
  • Distance From First - Distance in km/miles from now to oldest event.
  • Distance From Last - Distance in km/miles from now to newest event.
  • Distance Between - Distance between multiple events in km/miles into common statistics (min/max/mean/stddev).
  • Distance Between Points Within - Distance between 2 points, repeated across multiple events into common statistics (min/max/mean/stddev).

Other Features

  • Result Of Expression - Result of query language expression - expression may incorporate other features.
  • Enrichment Value - Looks up an enrichment value against a reference key, for the current value of the configured identifier.

Parameters of Features

Parameters of the features decide:

  • What calculation is performed
  • On what previous data; applying filter on previous data

How a Feature selects the events it computes over

Setting What it decides Values On
Feature Name The key the value is published under Text, unique per step All
Default Value What is published when a pivot attribute is missing from this event A number, or unset All
Scope How far the event search reaches
  • Same Node Instance
  • Same Organization
  • Same Primary Industry
  • Same Primary and Secondary Industry
  • Global
All but Result Of Expression
For Events Which event types count
  • All event types
  • Same as current
  • A named list of event types
All but the five with no event list
Include Current Event Whether this event counts towards its own Feature
  • On
  • Off, and off when unset
As For Events. Rejected outright by Approximate Distinct Count Of and Approximate Velocity Of
With The Same Which attribute values a past event has to share with this one 1 to 64 attributes, at least one an Identifier. Similarity 60 to 100 on Bitmap attributes As For Events. Approximate Velocity Of takes the expression form instead
Condition A further Query-Language filter on past events A query expression As For Events
Attributes The attributes the calculation itself runs over Identifiers and Subjects. Any attribute on the Approximates
  • Distinct Count Of
  • Specificity Of
  • Enrichment Value
  • The three Approximates
Time Window Which slice of time the window covers
  • all
  • mins
  • hours
  • days
  • weeks
  • months
  • events
  • since last event
  • since beginning of calendar day
  • since beginning of calendar month
All but Result Of Expression and Probability Of Time. Quantile Of Numeric, Specificity Of and the three Approximates take all and the fixed durations only
Starting Where the window ends, for excluding recent behaviour
  • immediately, and immediately when unset
  • mins in the past
  • hours in the past
  • days in the past
  • months in the past
  • events in the past
All but Quantile Of Numeric, Specificity Of, Result Of Expression and the three Approximates
Calculate Quantile Publishes the rank of this Feature's own result alongside it
  • On
  • Off, and off when unset
All
Into Features Names one Feature per statistic over the same criteria Any of the 17 statistic functions
  • Statistic Of Numeric
  • Statistic Of Expression
  • Time Between
  • Distance From
  • Distance Between
  • Distance Between Points Within
Coordinate Attribute The location a distance is measured from or between An attribute of Coordinate type The five Distance types
In Units What a distance is reported in
  • kilometers
  • miles
The five Distance types
Expression A Query-Language expression in place of an attribute A query expression
  • Statistic Of Expression
  • Result Of Expression
  • Approximate Distinct Count Of
  • Approximate Velocity Of
Numeric Attribute The number a statistic or quantile runs over An attribute of Real or Integer type
  • Statistic Of Numeric
  • Quantile Of Numeric
Normalize into Fraction of Total Whether a count is returned as its share of the population
  • On
  • Off
  • Specificity Of
  • Approximate Velocity Of
Reference Key Which enrichment the value is looked up in Text Enrichment Value
Bucket and Periodicity The short and the long window being compared
  • hours
  • days
  • weeks
  • months
Probability Of Time
Exclude Days left out of the calculation
  • weekdays
  • weekends
Probability Of Time

Default Value

Features perform calculations on attributes from both the current and past events. A feature may be configured to use a certain attribute, say profiling.device.identifier. In this case, if profiling.device.identifier is not populated within the current event then the feature cannot be computed. A default value can be optionally assigned in this case to make downstream rules, features or analysis easier.

The Default Value covers one case only: an attribute missing from the current event. It is written when any attribute the Feature pivots on is absent, which means the identifiers and subjects under With The Same, plus the attribute an Enrichment Value looks up. Where a Feature publishes several named outputs through Into Features, every one of them receives the Default Value.

It is not applied when the criteria simply matched nothing. If the attributes are present on the current event but no past event matches, the calculation runs over an empty set and publishes the result of that, which is usually zero. See Empty and single-event criteria.

With no Default Value set and an attribute missing, nothing is published at all: the Feature Name is absent from the map rather than present with a value.

Scope

Can be used to extend the event search from the local node, to all nodes within the same organization all the way to all across customers (global).

  • Same Node Instance - the event for the Feature will be sourced only from the node where the event is currently being processed. In cases where Production, Staging and Sandbox nodes exist, this option will prevent test data that is sent to the Staging and Sandbox environments from interfering and causing false positives in the Production node.
  • Same Organization - the events for the Feature will be sourced from all nodes within the current Organization. If Production, Staging and Sandbox nodes exist, then all of these nodes will be searched for matching events.
  • Global - the events for the Feature will be sourced across all Darwinium customers. This can be useful for checking if an Email Address, Device or other Identifier has been seen by another Darwinium customer to determine the age and other behavioural data.
    • Confirmation: The feature engine only ever has anonymized form of PII attributes to compute over, and only ever delivers the feature values as outputs (eg. velocity count, time since first).
  • Same Primary Industry - the events for the Feature will be sourced across all Darwinium customers with the same Primary Industry type. This can be used to eliminate high level conflicts within use cases - for example if a user has performed fraud in online retail, they may not be considered a bad actor in Health Care or Insurance.
  • Same Primary and Secondary Industry - a more fine grained industry filter using the Secondary Industry type to eliminate cross use-case false positives.

Set Scope explicitly. A Feature created in the Feature Editor is written as Same Node Instance. A Feature file that leaves Scope out is read as Same Organization, so the two routes do not agree and the difference is invisible on the page.

Scope is not offered on Result Of Expression, which reads other Features rather than events.

For Events

For Events provides a way of filtering past events based on their type. One of three settings applies.

  • All event Types - is a wildcard matching any event type. This is the default.
  • Same as current - matches events that only have the same type as the current event.
  • Specific event types - provides a user configurable list of event type matches. The types are listed under Event Type Reference on Creating and Editing Steps.

Include Current Event sits alongside them and decides whether the event being processed counts towards its own Feature. It is off by default, so a Velocity Of over a device counts past events only until it is turned on.

Usually better to leave as False

Including current event always includes the current event.

That can have unintended consequences such as considering the current event in scope even if the event_type doesn't match feature configuration or if there is a starting window.

It is usually safer to leave as False and handle the implications of only looking at prior events in the rule.

With The Same

Is used to specify which attributes from the current event must be present and match the values in the current event when a Feature is being calculated. This is somewhat analogous to creating a pivot table for the configured attributes.

At least one of them has to be an Identifier, and no more than 64 attributes may be listed. A Feature whose With The Same holds only Subjects is refused at build time.

Similarity turns an exact match into a fuzzy one, and appears only against attributes that support it: those with the Bitmap data type, plus profiling.device.identifier. It is a percentage and 100 (exact) is the default. A device signature that has drifted slightly still matches at 90, where it would not match at 100.

With The Same (expression) replaces the attribute list with a Query Language expression, for cases where the pivot is computed rather than named. Only the Approximate Feature Types offer it, and a Feature gives either the attribute list or the expression, never both.

Attributes

Some Feature types, such as Distinct Count Of perform a function on values for a configurable set of attributes. For such Features, the Attributes setting is where this configuration takes place.

For these Feature types, "With The Same" is somewhat analogous to creating a pivot table upon the selected values; In this analogy the Attributes setting is the equivalent of filtering the rows of the pivot table for the configured Attributes values.

Most Features with an Attributes configuration allow the use of only Subject and Identifier attributes.

Identifiers and Subjects are special attributes that are stored and accessible in real-time by Darwinium Features. Most Features require at least one Identifier to be specified. Refer to the specific documentation for a particular Feature for more detailed information.

A few Feature Types take exactly one attribute rather than a set: Specificity Of, Approximate Total Count Of and Enrichment Value.

Condition

Condition is an (optional) Query-Language filter that can be used to refine Features for specific use cases.

The following Condition could be used on a Velocity Of Feature to determine the number of high value purchases.

purchase.amount >= 10000

Time Window

Most Features have the option to specify a Time Window filter that can be used to restrict the input events to a fixed time period such as 6 hours, 5 days, 2 weeks etc.

Example: Time Window of 1month, Starting 7 days in the past. Notice how the Time Window is respected; the starting has shifts it backwards.Image

  • all - no window of the Feature's own. This is the default.
  • mins, hours, days, weeks, months - a fixed duration, counted back from the end of the window.
  • events - the last N matching events rather than a duration, so a quiet identifier reaches further back in tim Month is 30days.e than a busy one does.
  • since last event - bounds the window at the previous event rather than at a duration. Takes all event types, the same type as the current event, or a named list.
  • since beginning of calendar day, since beginning of calendar month - bounds the window at midnight, or at the first of the month, in a timezone picked from a list of UTC offsets. These are the two options that express a daily or a monthly limit.

Starting

For Features that have a Time Window option, the starting point of configured time window can be Immediate (meaning at the time the event occurs) or can be extended further back into the past, as a means of excluding recent behaviour.

  • immediately - the window ends at the current event. This is the default.
  • mins in the past, hours in the past, days in the past, months in the past - the window ends that far back instead, and the Time Window is measured from there. There is no weeks option; use 7, 14 or 21 days.
  • events in the past - skips that many of the most recent matching events, then measures the Time Window back from the one it lands on. The skip is capped at 32 events.

Example: Time Window of 1month, Starting 7 days in the past. Notice how the Time Window is respected; the starting has shifted it backwards.
Image

Starting is not offered on Result Of Expression, Approximate Distinct Count Of, Approximate Total Count Of, Quantile Of Numeric or Specificity Of. Approximate Velocity Of accepts immediately only.

Calculate Quantile

Any Feature can publish the quantile of its own result alongside the result itself. It is a single switch under Additional Settings in the Feature Editor, available on every Feature Type, and it is off by default.

With it on, the Feature's value is ranked against the distribution of that same Feature across the events this node has seen, and the rank is published to:

outcome['CHAMPION'].quantiles.general['<feature name>']

The raw value still goes to outcome['CHAMPION'].features.general['<feature name>'], so both are available to Rules, Models and Investigations.

This is how a threshold stops being a number somebody has to maintain. feature('logins_per_hour') >= 12 needs revisiting as traffic grows. The quantile form keeps meaning "busier than all but one percent of events" whatever the underlying counts do:

outcome['CHAMPION'].quantiles.general['logins_per_hour'] >= 0.99

Where a Feature produces several named outputs through Into Features, each output is ranked separately and published under its own name.

Quantile Of Numeric is already a quantile

Quantile Of Numeric ranks an attribute against the population as its whole purpose, and Probability Of Time publishes its own probability. Turning this on for either gives the quantile of a quantile, which is rarely what is wanted.



The settings below are offered only by some Feature Types. Each says which.

Into Features

Statistics functions are useful for behaviour and risk assessment, often multiple are required on the same criteria. To avoid the overhead of creating one Feature per function, the Into Features configuration is used to define criteria once and simply specify the Feature Name for each function. If a function is not required, simply leave the Feature Name blank.

Valid on: Statistic Of Numeric, Statistic Of Expression, Time Between, Distance From, Distance Between and Distance Between Points Within.

Statistic Functions

Each named function publishes one Feature under the name given to it. They are all computed over the same set of values: the numeric attribute for Statistic Of Numeric, the expression result for Statistic Of Expression, the gap between consecutive events for Time Between, and the distances for the Distance types.

is the number of values, each value and their mean.

Function Returns How it is calculated
Count Number of values
Sum
Minimum Lowest value
Maximum Greatest value
Range Max - Min
Average
Median Middle value of the sorted values Not interpolated. The value at position of the sorted list, counting from zero, so an even count returns the upper of the two middle values. Four values of 10, 20, 30, 40 give 30, not 25
Mode Most frequent value A tie publishes nothing. Where two or more values share the highest count, the Feature is left out of the map entirely rather than one of them being picked. A reader has to handle the name being absent
Median Absolute Deviation The median of the distances each value is from the median. A spread measure that is not pulled about by outliers the way Standard Deviation is. Both medians use the upper-middle rule above
Average Deviation The mean of the distances each value is from the mean, so a different statistic from Median Absolute Deviation rather than another name for it. Never publishes, see below
Variance Divides by , so this is the population variance, not the sample one
Standard Deviation
Sample Variance A more conservative spread estimate, knowing only a sample is being evaluated. Returns 0 rather than failing when there is a single value
Sample Standard Deviation
Coefficient Of Variation Undefined where the average is 0, which a Feature over values that are mostly zero can reach
Z Score The reference value is the oldest matching event in the window, not the current event. Returns 0 where every value is identical. Assuming a normal distribution, 0 is at the average, 0.5 the top 30%, 1 the top 16% and 2 the top 2%, with symmetry for negatives
Quantile Where the current value falls within all the values, between 0 and 1 Ties count strictly below against strictly above, so a value shared by the whole set returns 0.5. Not published by Statistic Of Expression, which has no current value to rank. On Distance From every distance is already relative to the current event, so this carries no information

Empty and single-event criteria

The statistics are computed over whatever the criteria matched, including nothing at all. The Default Value does not cover this case, so the numbers below are what Rules and Models actually read.

No matching events.

  • Count, Sum, Average, Median, Median Absolute Deviation, Range, Minimum, Maximum, Variance, Standard Deviation, Sample Variance, Sample Standard Deviation and Coefficient Of Variation all publish 0.
  • Mode publishes nothing.
  • Quantile publishes 0.5.
  • Z Score publishes infinity, which is worth guarding against in any rule that compares it.

One matching event.

  • Every spread measure is 0: Range, Variance, Standard Deviation, Sample Variance, Sample Standard Deviation, Median Absolute Deviation and Z Score, with Coefficient Of Variation joining them unless the value itself is 0.
  • A Z Score of 0 here means "nothing to compare against", not "average".
  • Minimum, Maximum, Average and Median are that value.
  • Time Between over one event is the case to watch. There is no gap to measure, so every statistic is 0, and yet Count publishes 1 because it counts events rather than gaps. A rule reading Minimum Time Between as a bot signal will fire on it.

A useful guard is to publish Count alongside whichever statistic the rule reads, and require it:

feature('purchase_amount_count') >= 3 and feature('purchase_amount_stddev') > 500

Coordinate Attribute

The attribute holding the location a distance is measured from or between. It must have the Coordinate data type; anything else is refused at build time.

Distance Between Points Within measures between two locations on the same event, so it takes a pair under the Comparing setting rather than one.

Valid on: the five distance Feature Types - Distance From, Distance From First, Distance From Last, Distance Between, Distance Between Points Within.

In Units

Whether a distance is reported in kilometers or miles. Those are the only two values. It has no default in the file format and must be set; the Feature Editor writes kilometers into a new Feature.

Valid on: the same five distance Feature Types as Coordinate Attribute.

Expression

A Query-Language expression used in place of an attribute. What it produces depends on the Feature Type: a numeric result to take statistics over, a value to count the unique occurrences of, or a result published as the Feature itself.

Valid on: Statistic Of Expression and Result Of Expression, where an expression is the only input. Approximate Distinct Count Of and Approximate Velocity Of, where it is an alternative to Attributes and exactly one of the two must be given.

Numeric Attribute

The single numeric attribute a statistic or quantile is calculated over. Only Real and Integer data types qualify, and each Feature Type adds its own restriction on top.

Valid on: Statistic Of Numeric, where the attribute must also be an Identifier or Subject. Quantile Of Numeric, where it must also be continuous.

Normalize into Fraction of Total

Whether the Feature returns a count or that count's share of the population. Set it and the count for this value is divided by the total count of all values the attribute has taken in the same window, giving a number between 0 and 1. Leave it and the count itself is returned. Both come from the same sketch, so normalizing costs nothing extra.

Valid on: Specificity Of and Approximate Velocity Of.

Bucket and Periodicity

The two windows Probability Of Time compares: Bucket is the short span of current activity being assessed, Periodicity the longer span the established pattern is drawn from. An hour against four weeks asks how usual this hour's activity is for this entity.

Both take hours, days, weeks or months. Minutes are not accepted on either, despite being offered. Periodicity has to use a unit at least as large as Bucket's and cover a longer duration, so hours against days is valid and days against hours is refused at build time.

Valid on: Probability Of Time.

Exclude

Removes weekdays or weekends from the calculation, so an entity whose weekday pattern is nothing like its weekend one can be assessed against the right half of its own history. Those are the only two values, and leaving it unset uses every day.

Valid on: Probability Of Time.