Enterprise Middleware

Subashi Pro WAF Middleware

Protect your Kipchak API with configurable WAF rules, rate limiting, and IP intelligence.

Introduction

Note: This middleware is only available with Kipchak Enterprise.

The Subashi Pro middleware is a Web Application Firewall (WAF) for Kipchak APIs. It builds on the existing Subashi WAF and evaluates requests against whitelist, blacklist, and rate-limit rules to help block common attack patterns and abusive traffic.

In addition to this, it adds Geo and ASN based IP intelligence, allowing you to block traffic by countries or ASN ranges.

Key features include:

  • Rule-Based Filtering: Combine multiple conditions per rule using a simple operator system.
  • Rate Limiting: Apply per-IP or custom key limits with a pluggable cache backend.
  • IP Intelligence: Geo-country and ASN matching against bundled MMDB databases - no API key, no network call on the request path.

Installation

To install this middleware, you need to access the Enterprise Composer repository at https://php.pkgs.1x.ax.

If you have an enterprise licence, please contact your account representative for access.

Once you have access (see https://getcomposer.org/doc/articles/authentication-for-private-packages.md on how to configure access once you have credentials), install the middleware via composer by running:

composer require kipchak/middleware-subashi-pro

Configuration

The middleware reads its configuration from kipchak.subashi.pro (for example in a subashi.pro.php config file).

Global Settings

  • enabled (bool): Master switch for the WAF.
  • blocked_response_code (int): Default HTTP status code for blocked requests.
  • blocked_response_message (string): Default response message for blocked requests.

Client IP Resolution

Every rule that depends on who is calling - ip, geo_country, asn and per-IP rate limiting - is evaluated against the address resolved here, so this is worth configuring deliberately.

Forwarding headers are set by whatever sits in front of the application, and anyone on the internet can send one. Subashi Pro therefore only reads them when the direct peer is a trusted proxy. The defaults trust nothing and use the connecting address: unspoofable, but it will be your load balancer's address if you have one in front.

  • client_ip.sources (array): Sources tried in order; the first that yields a valid address wins. Each entry is remote_addr or header:<Name>. Defaults to ['remote_addr'].
  • client_ip.trusted_proxies (array): Addresses and CIDR ranges, IPv4 or IPv6, whose forwarding headers may be believed. Empty (the default) means forwarding headers are never read. The literal '*' trusts any peer.
  • client_ip.chain_position (string): first (default) or last. Which end of a comma-separated header chain to read.

Put forwarding headers ahead of remote_addr in sources, or they will never be reached:

'client_ip' => [
    'sources' => ['header:X-Forwarded-For', 'remote_addr'],
    'trusted_proxies' => ['10.0.0.0/8'],
    'chain_position' => 'first',
],

If the trusted proxy sends no such header, resolution falls through to the next source, so keeping remote_addr last is a sensible backstop.

Headers with more than one value

X-Forwarded-For is a chain, not a single value:

X-Forwarded-For: 203.0.113.9, 70.41.3.18, 150.172.238.178

The leftmost entry is the original client as reported by the first proxy, and each hop appends the peer it saw. Only the rightmost entry was written by your own infrastructure - everything to its left is a claim the previous hop passed along, and a client can seed the chain by sending the header itself.

  • chain_position: 'first' reads 203.0.113.9. Correct when your edge proxy overwrites the header rather than appending to it, which is what Cloudflare, Kong and most CDNs do.
  • chain_position: 'last' reads 150.172.238.178. Cannot be spoofed, but with more than one hop in front it is your own proxy's address rather than the client's.

Entries that are not valid addresses - unknown, obfuscated identifiers - are skipped rather than failing the whole lookup. Port suffixes (203.0.113.9:41234), bracketed IPv6 ([2001:db8::1]:443) and zone identifiers (fe80::1%eth0) are stripped, and anything that is still not an IP address is discarded.

Rate Limiting

Rate limiting has its own section below: Rate Limiting.

Geolocation and ASN

Country and ASN are resolved locally from the MMDB databases shipped in kipchak/data-geo-asn, which is installed as a dependency. There is no API key to configure and no network call on the request path.

Two databases back each lookup and are consulted in order. If the first holds no record for the client address the second is tried; if neither does, the value is null, the condition does not match, and the request is let through.

LookupOrderDatabaseProviderRefreshed
Country1stdbip-country.mmdbDB-IP IP to Country LiteMonthly
Country2ndiptoasn-country.mmdbiptoasn.comDaily
ASN1stiptoasn-asn.mmdbiptoasn.comDaily
ASN2nddbip-asn.mmdbDB-IP IP to ASN LiteMonthly

A weekly pipeline publishes a refreshed kipchak/data-geo-asn, so composer update keeps the data current.

  • geolocation.enabled (bool): Enable country and ASN lookups.
  • geolocation.memo_limit (int): Client addresses memoised per worker process. Defaults to 1000; set to 0 to disable.
  • geolocation.country_databases (array): Optional. Absolute paths to country databases, in the order to consult them. Defaults to those bundled.
  • geolocation.asn_databases (array): Optional. Absolute paths to ASN databases, in the order to consult them. Defaults to those bundled.

Under FrankenPHP worker mode the readers are opened once per worker and held, so the cost per request is a single in-memory trie walk. For very high lookup volumes, install the optional maxmind-db/reader-ext C extension; it is a drop-in replacement and needs no configuration change.

Results derived from the DB-IP databases carry a CC BY 4.0 attribution requirement. Blocking a request is not a display of the data, so nothing is required for firewall use, but if you surface country or ASN in an API response or UI you must credit IP geolocation by DB-IP.

Diagnostics

Getting client_ip right means knowing which headers actually reach the application and what your proxies put in them — which is precisely what you cannot see from outside. The diagnostics dump puts it on screen.

  • diagnostics.enabled (bool): Off by default.
  • diagnostics.token (string): Shared secret. Diagnostics stay inert while this is empty, even when enabled.
  • diagnostics.query_param (string): Query parameter carrying the token. Defaults to kipchak-waf-debug.
'diagnostics' => [
    'enabled' => true,
    'query_param' => 'kipchak-waf-debug',
    'token' => env('SUBASHI_DIAGNOSTICS_TOKEN', ''),
],

Add the parameter to any route and the middleware answers with the dump instead of passing the request on:

https://api.example.com/v1/anything?kipchak-waf-debug=your-token

It reports the resolved client address and which source produced it, whether the connecting peer counted as a trusted proxy, the country and ASN with which of the four databases answered, every header on the request, and a narrow slice of server parameters:

{
  "client_ip": {
    "resolved": "212.58.244.20",
    "resolved_from": "header:X-Forwarded-For",
    "remote_addr": "10.1.2.3",
    "peer_is_trusted_proxy": true,
    "configured_sources": ["header:X-Forwarded-For", "remote_addr"],
    "configured_trusted_proxies": ["10.0.0.0/8"],
    "chain_position": "first"
  },
  "geo": {
    "country": { "value": "GB", "database": "dbip-country.mmdb" },
    "asn": { "value": "2818", "organisation": "BBC Internet Services, UK", "database": "iptoasn-asn.mmdb" },
    "data_version": "2026.09.06"
  },
  "headers": { "...": "every header, repeated ones kept as a list" }
}

The check runs before every other rule, and before the enabled switch, so you can still ask "why does the firewall see me as this address?" on a request that would otherwise be blocked.

Two deliberate choices about what the dump contains. Headers are not redacted — they are the caller's own headers, and hiding the Authorization header from someone debugging their Authorization header would defeat the point. Server parameters, by contrast, are limited to a fixed list of network-related keys, because on most SAPIs that array carries the process environment and would otherwise hand over your database password.

That combination is why the token is mandatory rather than optional: the dump shows a caller their own credentials and describes your proxy topology. Each use is logged at warning level with the connecting address and path.

Logging

A rule's name does not affect matching — the list a rule sits in decides the outcome — but it is what identifies the rule afterwards, so make it something you would want to read in an alert.

EventLevelContext
Request blockedwarningrule, ip, method, path, user_agent, response_code
Request rate limitedwarningrule, ip, path, limit, window
Request whitelisteddebugrule, ip, path
Rate limit store unavailableerrorerror (at most once a minute per worker)

Blocks and throttles log at warning because they are the events you want visible without turning on debug logging. Whitelist hits are routine traffic and log at debug.

Rules

Rules are declared under whitelist, blacklist, and rate_limit_rules. Each rule has a name and a list of conditions. All conditions must match for the rule to apply.

Condition Types

  • header - HTTP header value
  • query_param - A single named URL query parameter
  • query - The whole query string, URL-decoded once. Use this to scan for patterns without knowing the parameter names in advance
  • ip - Client IP address, as resolved by the rules in Client IP Resolution
  • geo_country - ISO 3166-1 alpha-2 country code (from the bundled country databases)
  • asn - Autonomous System Number, as a string (from the bundled ASN databases)
  • method - HTTP method
  • path - URL path only, not the query string
  • body - JSON body field

Operators

  • equals
  • not_equals
  • contains
  • not_contains
  • regex
  • in_list
  • not_in_list
  • exists
  • not_exists
  • gt
  • lt
  • gte
  • lte

Rate Limiting

Rate limiting caps how many requests a client can make in a period, to protect the API from abuse: scrapers, credential stuffing, runaway scripts. It runs in the firewall, before authentication, so it can only tell clients apart by what the request itself carries: the IP address, or a secret such as an API key or a bearer token.

How requests are counted

Each rule allows limit requests in any window seconds. The count is a sliding window: a request is weighed against the requests in the current window plus the share of the previous window that still falls within the last window seconds. A client therefore cannot make limit requests at the end of one window and limit more at the start of the next.

Requests that are refused are not counted, so a client that keeps retrying while limited is let back in at the rate the limit allows.

Counters are shared by every worker, and their increments are atomic, so a limit holds when many requests arrive at once. In testing, 60 simultaneous requests against a limit of 20 let exactly 20 through, across four FrankenPHP workers, on both Memcached and Valkey.

storePackageScope
memcachedkipchak/driver-memcachedEvery host that shares the Memcached pool. Atomic increment and add.
valkeykipchak/driver-valkey (Enterprise)Every host that shares the Valkey pool. Atomic INCR in a Lua script.
filekipchak/driver-filecacheOne host only. Atomic between the workers of that host through file locks.

Use memcached or valkey when the API runs on more than one host: with file, each host counts separately, so the effective limit is multiplied by the number of hosts.

Configuration

'rate_limiting' => [
    'enabled' => true,
    'default_limit' => 100,      // requests per window, for rules that do not set their own
    'default_window' => 60,      // seconds
    'store' => 'memcached',      // 'memcached', 'valkey' or 'file'
    'memcached_pool' => 'cache', // when store is 'memcached'
    'valkey_pool' => 'cache',    // when store is 'valkey'
    'fail_open' => true,         // when the store is unreachable: true allows requests, false refuses them
    'headers' => false,          // add RateLimit-Policy and RateLimit headers to responses
],
KeyDefaultDescription
enabledfalseTurns rate limiting on.
default_limit60Requests per window for rules without rate_limit.limit.
default_window60Window in seconds for rules without rate_limit.window.
storefileWhere counters are kept.
memcached_poolcacheThe Memcached driver pool.
valkey_poolcacheThe Valkey driver pool.
fail_opentrueWhat happens when the store cannot be reached.
headersfalseWhether responses carry the rate limit headers.

Each rule in rate_limit_rules has a name, optional conditions, and a rate_limit:

'rate_limit_rules' => [
    [
        'name' => 'Per IP',
        'conditions' => [],               // no conditions: applies to every request
        'rate_limit' => [
            'limit' => 100,
            'window' => 60,
            'key_prefix' => 'ip',         // keeps this rule's counters apart from other rules'
            'key_source' => 'ip',         // what tells one client from another
        ],
    ],
],

key_source is one of:

key_sourceCounts per
ipClient IP address.
header:<Name>Value of a request header, e.g. header:apikey.
query:<name>Value of a query parameter, e.g. query:apikey.

Values are hashed before they are used as counter keys, so API keys and tokens never reach the store. A request without the header or parameter is counted under one shared unknown key for that rule.

Every rule whose conditions match counts the request against its own limit, and the first rule over its limit refuses it. Whitelisted requests are not counted at all.

Responses

A refused request receives 429 Too Many Requests with a Retry-After header giving the number of seconds until a request is likely to be allowed again:

HTTP/1.1 429 Too Many Requests
Retry-After: 18
X-Protected-By: Subashi Pro by Mamluk
Content-Type: application/json

{"code":429,"status":"TOO MANY REQUESTS","data":"Rate limit exceeded"}

With 'headers' => true, every response from a rate-limited route also carries the RateLimit-Policy and RateLimit fields of the IETF draft RateLimit header fields for HTTP, describing the matching rule with the fewest requests remaining:

RateLimit-Policy: "Per API key";q=1000;w=3600
RateLimit: "Per API key";r=742;t=1260

q is the limit, w the window in seconds, r the requests remaining and t the seconds until the current window ends. Well-behaved clients can slow down before they are refused.

When the store is down

With fail_open set to true (the default), requests are allowed while the store cannot be reached, so a cache outage does not become an outage of the API. With false, requests are refused with 503 Service Unavailable. Either way the failure is logged at error level, at most once a minute per worker, so an outage does not flood the log. Workers reconnect to Valkey by themselves when it returns.

Examples

One limit per IP address

[
    'name' => 'Per IP',
    'conditions' => [],
    'rate_limit' => ['limit' => 300, 'window' => 60, 'key_prefix' => 'ip', 'key_source' => 'ip'],
],

Per-IP limits use the address resolved by Client IP Resolution. Behind a load balancer or CDN, configure client_ip so that each client behind the proxy is counted separately; forwarding headers from untrusted peers are ignored, so a client cannot dodge its limit by sending a new X-Forwarded-For.

A stricter limit for one endpoint

Rules combine, so a login endpoint can have its own, much lower limit on top of the general one:

[
    'name' => 'Login attempts per IP',
    'conditions' => [
        ['type' => 'path', 'operator' => 'equals', 'value' => '/v1/login'],
        ['type' => 'method', 'operator' => 'equals', 'value' => 'POST'],
    ],
    'rate_limit' => ['limit' => 5, 'window' => 300, 'key_prefix' => 'login', 'key_source' => 'ip'],
],

A stricter limit for requests without credentials

[
    'name' => 'Unauthenticated per IP',
    'conditions' => [
        ['type' => 'header', 'key' => 'apikey', 'operator' => 'not_exists'],
        ['type' => 'header', 'key' => 'Authorization', 'operator' => 'not_exists'],
    ],
    'rate_limit' => ['limit' => 20, 'window' => 60, 'key_prefix' => 'anon', 'key_source' => 'ip'],
],

With the authentication middlewares

The firewall should run before authentication, so that abusive traffic is turned away before any key is looked up or token verified. In middlewares/middlewares.php Slim runs the middleware added last first, so initialise Subashi after the authentication middlewares:

Error::initialise($app);
Key::initialise($app);      // or JWKS, JWT, HMAC
SubashiPro::initialise($app);  // added last, so it runs first

At that point nothing has been verified. Key rate limits only on the IP address or on a secret the client holds. Never key on a public identifier, such as a client ID, an HMAC key ID or a tenant header: anyone can send someone else's identifier, and would use up that client's allowance.

API keys (auth-key)

With API key authentication, clients are defined in kipchak.auth.key, and each request carries its key in the apikey header or query parameter (or whichever you configured there). The key is a secret, so it identifies the client safely:

// config/kipchak.auth.key.php
'authorised_keys' => [
    'k_live_8f2c...' => 'acme-corp',
    'k_live_31d0...' => 'globex',
],
'header' => 'apikey',
'query_param' => 'apikey',
// config/kipchak.subashi.pro.php
'rate_limit_rules' => [
    [
        'name' => 'Per API key (header)',
        'conditions' => [['type' => 'header', 'key' => 'apikey', 'operator' => 'exists']],
        'rate_limit' => ['limit' => 1000, 'window' => 3600, 'key_prefix' => 'apikey', 'key_source' => 'header:apikey'],
    ],
    [
        'name' => 'Per API key (query)',
        'conditions' => [
            ['type' => 'header', 'key' => 'apikey', 'operator' => 'not_exists'],
            ['type' => 'query_param', 'key' => 'apikey', 'operator' => 'exists'],
        ],
        'rate_limit' => ['limit' => 1000, 'window' => 3600, 'key_prefix' => 'apikey', 'key_source' => 'query:apikey'],
    ],
    [
        // Invalid keys get their own counter each, so cap every caller by address as well.
        'name' => 'Per IP',
        'conditions' => [],
        'rate_limit' => ['limit' => 300, 'window' => 60, 'key_prefix' => 'ip', 'key_source' => 'ip'],
    ],
],

Both API key rules share the prefix apikey, so a client is counted once whether it sends its key in the header or the query string.

To give some clients a higher limit, match their keys with in_list and exclude them from the general rule with not_in_list. Every matching rule counts, so without the exclusion they would also be held to the lower limit:

[
    'name' => 'Partners',
    'conditions' => [['type' => 'header', 'key' => 'apikey', 'operator' => 'in_list', 'values' => [env('ACME_API_KEY', '')]]],
    'rate_limit' => ['limit' => 10000, 'window' => 3600, 'key_prefix' => 'partner', 'key_source' => 'header:apikey'],
],
[
    'name' => 'Everyone else',
    'conditions' => [
        ['type' => 'header', 'key' => 'apikey', 'operator' => 'exists'],
        ['type' => 'header', 'key' => 'apikey', 'operator' => 'not_in_list', 'values' => [env('ACME_API_KEY', '')]],
    ],
    'rate_limit' => ['limit' => 1000, 'window' => 3600, 'key_prefix' => 'apikey', 'key_source' => 'header:apikey'],
],

This puts the keys in two config files. Limits per client and plan, by the client name in kipchak.auth.key, are what Quotas are for.

Tokens from an identity provider (auth-jwks, auth-jwt)

With JWKS or JWT authentication, clients are not defined in Kipchak: your identity provider issues signed tokens, and the client's identity is a claim inside them, such as sub, client_id or azp. The firewall runs before the token is verified, so it must not read claims: an unverified token can claim to be anyone.

The bearer token itself is a secret, so it is safe to count per token:

[
    'name' => 'Per bearer token',
    'conditions' => [['type' => 'header', 'key' => 'Authorization', 'operator' => 'exists']],
    'rate_limit' => ['limit' => 600, 'window' => 60, 'key_prefix' => 'bearer', 'key_source' => 'header:Authorization'],
],
[
    'name' => 'Per IP',
    'conditions' => [],
    'rate_limit' => ['limit' => 300, 'window' => 60, 'key_prefix' => 'ip', 'key_source' => 'ip'],
],

Each new access token starts a new count, so this limits bursts within a token's lifetime rather than a client's total use. Limits per client or per plan, keyed by the verified sub or client_id, need to run after the token is verified. That is what Quotas are for.

Signed requests (auth-hmac)

Request Signing clients send a new Kipchak-Signature on every request, and the key ID inside it is not a secret, so neither identifies a client safely before the signature is verified. Limit signed traffic per IP in the firewall:

[
    'name' => 'Signed API calls per IP',
    'conditions' => [['type' => 'header', 'key' => 'Kipchak-Signature', 'operator' => 'exists']],
    'rate_limit' => ['limit' => 1200, 'window' => 60, 'key_prefix' => 'signed', 'key_source' => 'ip'],
],

Webhook senders such as Stripe deliver from many addresses in bursts. Whitelist their paths, or give them a generous rule of their own, rather than holding them to a per-IP limit meant for API clients.

Quotas

Per-plan quotas, added to Subashi Pro in 1.31, are their own package from Subashi Pro 1.32: Quotas. An API can now enforce plans without the firewall. To upgrade, see Upgrading from Subashi Pro 1.31.

Usage

Once enabled, the middleware evaluates requests in this order:

  1. Whitelist: Matching rules bypass all other checks.
  2. Blacklist: Matching rules block the request (with optional custom response).
  3. Rate limiting: Matching rules enforce request limits.

Example Blacklist Rule

[
    'name' => 'Block SQL injection attempts',
    'conditions' => [
        [
            'type' => 'path',
            'operator' => 'regex',
            'value' => '/(union.*select|select.*from|drop.*table)/i',
        ],
    ],
    'response_code' => 403,
    'response_message' => 'SQL injection attempt detected',
]

Example Geo Country Rule

[
    'name' => 'Block specific countries',
    'conditions' => [
        [
            'type' => 'geo_country',
            'operator' => 'in_list',
            'values' => ['RU', 'CN'],
        ],
    ],
]

Example ASN Rule

[
    'name' => 'Block known ASN ranges',
    'conditions' => [
        [
            'type' => 'asn',
            'operator' => 'in_list',
            'values' => ['13335', '15169'],
        ],
    ],
]

Rate limit examples are in Rate Limiting.

Git Repository

The source code for this middleware is hosted internally.

Previous
Auth - JWT