Saturday, 18 April 2020

Testing http/2 on Apache and Centos 8

Back in 2015 Apache (2.4.17 and onwards) started to support http/2 protocol through a dedicated apache module mod_http2.
CentOS 8  ships with Apache/2.4.37 and thus is cabable of serving http/2 and it was worth a test to try to see how this innovative protocol update works.

Apache on CentOS 8 ships with mod_http2 enabled by default, but in order to use the http/2 protocol, one needs to specify it expelicity using Protocols directive. (Note that it is Protocols with an 's').
Below is a sample Apache config that enables both http/2 on plain/clear text 'h2c' and standard h2 which works on top of SSL using SSL ALPN(Application Layer Protocol Negociation).
Please note that SSL must be enabled thus Openssl should be installed and Apache mod_ssl should also be installed and an https should be configured for h2 to work.


Apache config:

[root@beren ~]# cat /etc/httpd/conf.d/http2link.conf
### Adding http2 link headers and  H2PushResource

Protocols h2c h2 http/1.1
H2EarlyHints on

Header add Link "</test/ysf_100.png>; rel=preload; as=image"
Header add Link "</test/ysf_099.png>; rel=preload; as=image"
Header add Link "</test/ysf_098.png>; rel=preload; as=image"
Header add Link "</test/ysf_097.png>; rel=preload; as=image"
Header add Link "</test/ysf_096.png>; rel=preload; as=image"

H2PushResource /test/ysf.png
H2PushResource /test/ysf_096.png
H2PushResource /test/ysf_095.png
H2PushResource /test/ysf_094.png
H2PushResource /test/ysf_093.png
H2PushResource /test/ysf_092.png

[root@beren ~]#


The configuration mainly contains the Protocols directive which lists the prefered protocls Apache will offer to the client starting with prefered from left to right, ordering matters as per Apache documentation.

Then we expelicty set early hints to on, this feature will make use of the http/2 server push features to speed up page load times.
There are 2 ways Apache can use the http/2 server push, either by adding the 'Link' header as shown above using mod_header 'Header add' directive or using the new mod_http2 'H2PushResource' directive to push a certain resource to the client using early hints.
I have created an HTML page with some 110 image resources and tried to push the above subset of resources to test the configuration.

I tried 3 different clients, curl (curl now supports http/2 if compiled with the nghttp2 library, CentOS 8 offer curl compiled with http/2 feature), nghttp client tool and Vivaldi browser (Chromium like).

Below is the output of curl in connecting to Apache using h2c (http/2 on clear text), the protocol upgrade is visible and you can see indeed there is an extra reponse with HTTP 101 Protocol upgrade before the standard HTTP 200 response:

sherif@fingolfin:~$ curl -v --http2 http://192.168.56.105/test/test.html >/dev/null
*   Trying 192.168.56.105...
* TCP_NODELAY set
  % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current
                                 Dload  Upload   Total   Spent    Left  Speed
  0     0    0     0    0     0      0      0 --:--:-- --:--:-- --:--:--     0* Connected to 192.168.56.105 (192.168.56.105) port 80 (#0)
> GET /test/test.html HTTP/1.1
> Host: 192.168.56.105
> User-Agent: curl/7.58.0
> Accept: */*
> Connection: Upgrade, HTTP2-Settings
> Upgrade: h2c
> HTTP2-Settings: AAMAAABkAARAAAAAAAIAAAAA
>
< HTTP/1.1 101 Switching Protocols
< Upgrade: h2c
< Connection: Upgrade
* Received 101
* Using HTTP2, server supports multi-use
* Connection state changed (HTTP/2 confirmed)
* Copying HTTP/2 data in stream buffer to connection buffer after upgrade: len=0
* Connection state changed (MAX_CONCURRENT_STREAMS updated)!
< HTTP/2 103
< link: </test/ysf.png>; rel=preload, </test/ysf_096.png>; rel=preload, </test/ysf_095.png>; rel=preload, </test/ysf_094.png>; rel=preload, </test/ysf_093.png>; rel=preload, </test/ysf_092.png>; rel=preload
< HTTP/2 200
< date: Sun, 00 Jan 1900 00:00:00 GMT
< server: Apache/2.4.37 (centos) OpenSSL/1.1.1c
< last-modified: Fri, 17 Apr 2020 18:57:51 GMT
< etag: W/"1326-5a3812047b27b"
< accept-ranges: bytes
< content-length: 4902
< link: </test/ysf_100.png>; rel=preload; as=image
< link: </test/ysf_099.png>; rel=preload; as=image
< link: </test/ysf_098.png>; rel=preload; as=image
< link: </test/ysf_097.png>; rel=preload; as=image
< link: </test/ysf_096.png>; rel=preload; as=image
< content-type: text/html; charset=UTF-8
<
{ [4902 bytes data]
100  4902  100  4902    0     0  2393k      0 --:--:-- --:--:-- --:--:-- 2393k
* Connection #0 to host 192.168.56.105 left intact
sherif@fingolfin:~$ 

You can also see the Link headers added to the response from Apache with preload, this should direct the browser to receive those resoruces early on to speed up page loading.
Then below is the output of using http/2 over SSL, this time using h2 with SSL ALPN negociation, Apache is seen offering both http/2 then http/1.1 using ALPN and using TLS1.3 during the handshake:

sherif@fingolfin:~$ curl -v -k --http2 https://192.168.56.105/test/test.html >/dev/null
*   Trying 192.168.56.105...
* TCP_NODELAY set
  % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current
                                 Dload  Upload   Total   Spent    Left  Speed
  0     0    0     0    0     0      0      0 --:--:-- --:--:-- --:--:--     0* Connected to 192.168.56.105 (192.168.56.105) port 443 (#0)
* ALPN, offering h2
* ALPN, offering http/1.1
* successfully set certificate verify locations:
*   CAfile: /etc/ssl/certs/ca-certificates.crt
  CApath: /etc/ssl/certs
} [5 bytes data]
* TLSv1.3 (OUT), TLS handshake, Client hello (1):
} [512 bytes data]
* TLSv1.3 (IN), TLS handshake, Server hello (2):
{ [122 bytes data]
* TLSv1.3 (IN), TLS Unknown, Certificate Status (22):
{ [1 bytes data]
* TLSv1.3 (IN), TLS handshake, Unknown (8):
{ [15 bytes data]
* TLSv1.3 (IN), TLS Unknown, Certificate Status (22):
{ [1 bytes data]
* TLSv1.3 (IN), TLS handshake, Certificate (11):
{ [2672 bytes data]
* TLSv1.3 (IN), TLS Unknown, Certificate Status (22):
{ [1 bytes data]
* TLSv1.3 (IN), TLS handshake, CERT verify (15):
{ [264 bytes data]
* TLSv1.3 (IN), TLS Unknown, Certificate Status (22):
{ [1 bytes data]
* TLSv1.3 (IN), TLS handshake, Finished (20):
{ [52 bytes data]
* TLSv1.3 (OUT), TLS change cipher, Client hello (1):
} [1 bytes data]
* TLSv1.3 (OUT), TLS Unknown, Certificate Status (22):
} [1 bytes data]
* TLSv1.3 (OUT), TLS handshake, Finished (20):
} [52 bytes data]
* SSL connection using TLSv1.3 / TLS_AES_256_GCM_SHA384
* ALPN, server accepted to use h2
* Server certificate:
*  subject: C=US; O=Unspecified; CN=beren; emailAddress=root@beren
*  start date: Apr 16 19:10:23 2020 GMT
*  expire date: Apr 21 20:50:23 2021 GMT
*  issuer: C=US; O=Unspecified; OU=ca-5235634170413358813; CN=beren; emailAddress=root@beren
*  SSL certificate verify result: self signed certificate in certificate chain (19), continuing anyway.
* Using HTTP2, server supports multi-use
* Connection state changed (HTTP/2 confirmed)
* Copying HTTP/2 data in stream buffer to connection buffer after upgrade: len=0
} [5 bytes data]
* TLSv1.3 (OUT), TLS Unknown, Unknown (23):
} [1 bytes data]
* TLSv1.3 (OUT), TLS Unknown, Unknown (23):
} [1 bytes data]
* TLSv1.3 (OUT), TLS Unknown, Unknown (23):
} [1 bytes data]
* Using Stream ID: 1 (easy handle 0x55e221b8d580)
} [5 bytes data]
* TLSv1.3 (OUT), TLS Unknown, Unknown (23):
} [1 bytes data]
> GET /test/test.html HTTP/2
> Host: 192.168.56.105
> User-Agent: curl/7.58.0
> Accept: */*
>
{ [5 bytes data]
* TLSv1.3 (IN), TLS Unknown, Certificate Status (22):
{ [1 bytes data]
* TLSv1.3 (IN), TLS handshake, Newsession Ticket (4):
{ [265 bytes data]
* TLSv1.3 (IN), TLS Unknown, Certificate Status (22):
{ [1 bytes data]
* TLSv1.3 (IN), TLS handshake, Newsession Ticket (4):
{ [265 bytes data]
* TLSv1.3 (IN), TLS Unknown, Unknown (23):
{ [1 bytes data]
* Connection state changed (MAX_CONCURRENT_STREAMS updated)!
} [5 bytes data]
* TLSv1.3 (OUT), TLS Unknown, Unknown (23):
} [1 bytes data]
* TLSv1.3 (IN), TLS Unknown, Unknown (23):
{ [1 bytes data]
* TLSv1.3 (IN), TLS Unknown, Unknown (23):
{ [1 bytes data]
< HTTP/2 103
< link: </test/ysf.png>; rel=preload, </test/ysf_096.png>; rel=preload, </test/ysf_095.png>; rel=preload, </test/ysf_094.png>; rel=preload, </test/ysf_093.png>; rel=preload, </test/ysf_092.png>; rel=preload
{ [5 bytes data]
* TLSv1.3 (IN), TLS Unknown, Unknown (23):
{ [1 bytes data]
< HTTP/2 200
< date: Fri, 17 Apr 2020 19:59:34 GMT
< server: Apache/2.4.37 (centos) OpenSSL/1.1.1c
< last-modified: Fri, 17 Apr 2020 18:57:51 GMT
< etag: "1326-5a3812047b27b"
< accept-ranges: bytes
< content-length: 4902
< link: </test/ysf_100.png>; rel=preload; as=image
< link: </test/ysf_099.png>; rel=preload; as=image
< link: </test/ysf_098.png>; rel=preload; as=image
< link: </test/ysf_097.png>; rel=preload; as=image
< link: </test/ysf_096.png>; rel=preload; as=image
< content-type: text/html; charset=UTF-8
<
{ [978 bytes data]
* TLSv1.3 (IN), TLS Unknown, Unknown (23):
{ [1 bytes data]
* TLSv1.3 (IN), TLS Unknown, Unknown (23):
{ [1 bytes data]
* TLSv1.3 (IN), TLS Unknown, Unknown (23):
{ [1 bytes data]
* TLSv1.3 (IN), TLS Unknown, Unknown (23):
{ [1 bytes data]
100  4902  100  4902    0     0  76593      0 --:--:-- --:--:-- --:--:-- 76593
* Connection #0 to host 192.168.56.105 left intact
sherif@fingolfin:~$

The next test is done with the nghttp client tool that ships with nghttp2/1.33.0 and is a client implementation for the http/2 C liberary libnghttp2.
The main objective here is to test resoruces being pushed and seening if it works as expected and the differance between resources pushed with early hints and resources pushed with Link header:

Testing pushes:

sherif@fingolfin:~$ nghttp -vnys https://192.168.56.105/test/test.html
[  0.003] Connected
The negotiated protocol: h2
[  0.054] recv SETTINGS frame <length=6, flags=0x00, stream_id=0>
          (niv=1)
          [SETTINGS_MAX_CONCURRENT_STREAMS(0x03):100]
[  0.054] recv WINDOW_UPDATE frame <length=4, flags=0x00, stream_id=0>
          (window_size_increment=2147418112)
[  0.054] send SETTINGS frame <length=12, flags=0x00, stream_id=0>
          (niv=2)
          [SETTINGS_MAX_CONCURRENT_STREAMS(0x03):100]
          [SETTINGS_INITIAL_WINDOW_SIZE(0x04):65535]
[  0.054] send SETTINGS frame <length=0, flags=0x01, stream_id=0>
          ; ACK
          (niv=0)
[  0.054] send PRIORITY frame <length=5, flags=0x00, stream_id=3>
          (dep_stream_id=0, weight=201, exclusive=0)
[  0.054] send PRIORITY frame <length=5, flags=0x00, stream_id=5>
          (dep_stream_id=0, weight=101, exclusive=0)
[  0.054] send PRIORITY frame <length=5, flags=0x00, stream_id=7>
          (dep_stream_id=0, weight=1, exclusive=0)
[  0.054] send PRIORITY frame <length=5, flags=0x00, stream_id=9>
          (dep_stream_id=7, weight=1, exclusive=0)
[  0.054] send PRIORITY frame <length=5, flags=0x00, stream_id=11>
          (dep_stream_id=3, weight=1, exclusive=0)
[  0.054] send HEADERS frame <length=50, flags=0x25, stream_id=13>
          ; END_STREAM | END_HEADERS | PRIORITY
          (padlen=0, dep_stream_id=11, weight=16, exclusive=0)
          ; Open new stream
          :method: GET
          :path: /test/test.html
          :scheme: https
          :authority: 192.168.56.105
          accept: */*
          accept-encoding: gzip, deflate
          user-agent: nghttp2/1.30.0
[  0.056] recv SETTINGS frame <length=0, flags=0x01, stream_id=0>
          ; ACK
          (niv=0)
[  0.057] recv (stream_id=13) :scheme: https
[  0.057] recv (stream_id=13) :authority: 192.168.56.105
[  0.057] recv (stream_id=13) :path: /test/ysf.png
[  0.057] recv (stream_id=13) :method: GET
[  0.057] recv (stream_id=13) accept: */*
[  0.057] recv (stream_id=13) accept-encoding: gzip, deflate
[  0.057] recv (stream_id=13) user-agent: nghttp2/1.30.0
[  0.057] recv (stream_id=13) host: 192.168.56.105
[  0.057] recv PUSH_PROMISE frame <length=60, flags=0x04, stream_id=13>
          ; END_HEADERS
......          (padlen=0)
          ; First push response header
[  0.067] recv (stream_id=20) :status: 200
[  0.067] recv (stream_id=20) date: Fri, 17 Apr 2020 20:02:08 GMT
[  0.067] recv (stream_id=20) server: Apache/2.4.37 (centos) OpenSSL/1.1.1c
[  0.067] recv (stream_id=20) last-modified: Thu, 16 Apr 2020 20:38:41 GMT
[  0.067] recv (stream_id=20) etag: "36ec-5a36e6b16758e"
[  0.067] recv (stream_id=20) accept-ranges: bytes
[  0.067] recv (stream_id=20) content-length: 14060
[  0.067] recv (stream_id=20) link: </test/ysf_100.png>; rel=preload; as=image
[  0.067] recv (stream_id=20) link: </test/ysf_099.png>; rel=preload; as=image
[  0.067] recv (stream_id=20) link: </test/ysf_098.png>; rel=preload; as=image
[  0.067] recv (stream_id=20) link: </test/ysf_097.png>; rel=preload; as=image
[  0.067] recv (stream_id=20) link: </test/ysf_096.png>; rel=preload; as=image
[  0.067] recv (stream_id=20) content-type: image/png
[  0.067] recv HEADERS frame <length=37, flags=0x04, stream_id=20>
          ; END_HEADERS
          (padlen=0)
[  0.067] recv (stream_id=2) :status: 103
[  0.067] recv (stream_id=2) link: </test/ysf.png>; rel=preload, </test/ysf_096.png>; rel=preload, </test/ysf_095.png>; rel=preload, </test/ysf_094.png>; rel=preload, </test/ysf_093.png>; rel=preload, </test/ysf_092.png>; rel=preload
[  0.067] recv HEADERS frame <length=2, flags=0x04, stream_id=2>
          ; END_HEADERS
          (padlen=0)
          ; First push response header
[  0.068] recv (stream_id=2) :status: 200
[  0.068] recv (stream_id=2) date: Fri, 17 Apr 2020 20:02:08 GMT
[  0.068] recv (stream_id=2) server: Apache/2.4.37 (centos) OpenSSL/1.1.1c
[  0.068] recv (stream_id=2) last-modified: Thu, 16 Apr 2020 20:37:22 GMT
[  0.068] recv (stream_id=2) etag: "36ec-5a36e665c587d"
........[  0.079] recv DATA frame <length=1291, flags=0x00, stream_id=20>
[  0.079] recv DATA frame <length=487, flags=0x01, stream_id=14>
          ; END_STREAM
[  0.079] recv DATA frame <length=1291, flags=0x00, stream_id=2>
[  0.079] recv DATA frame <length=1150, flags=0x01, stream_id=20>
          ; END_STREAM
[  0.079] recv DATA frame <length=167, flags=0x01, stream_id=2>
          ; END_STREAM
[  0.079] send GOAWAY frame <length=8, flags=0x00, stream_id=0>
          (last_stream_id=20, error_code=NO_ERROR(0x00), opaque_data(0)=[])
***** Statistics *****

Request timing:
  responseEnd: the  time  when  last  byte of  response  was  received
               relative to connectEnd
 requestStart: the time  just before  first byte  of request  was sent
               relative  to connectEnd.   If  '*' is  shown, this  was
               pushed by server.
      process: responseEnd - requestStart
         code: HTTP status code
         size: number  of  bytes  received as  response  body  without
               inflation.
          URI: request URI

see http://www.w3.org/TR/resource-timing/#processing-model

sorted by 'complete'

id  responseEnd requestStart  process code size request path
 13     +9.49ms       +675us   8.81ms  200   4K /test/test.html
  4    +13.76ms *    +3.93ms   9.83ms  200  13K /test/ysf_096.png
  6    +23.55ms *    +4.11ms  19.44ms  200  13K /test/ysf_095.png
  8    +23.60ms *    +4.29ms  19.31ms  200  13K /test/ysf_094.png
 10    +23.65ms *    +4.68ms  18.96ms  200  13K /test/ysf_093.png
 12    +23.93ms *    +6.61ms  17.32ms  200  13K /test/ysf_092.png
 16    +25.72ms *    +8.12ms  17.59ms  200  13K /test/ysf_099.png
 18    +25.75ms *    +8.29ms  17.46ms  200  13K /test/ysf_098.png
 14    +25.81ms *    +7.96ms  17.85ms  200  13K /test/ysf_100.png
 20    +25.87ms *    +8.58ms  17.28ms  200  13K /test/ysf_097.png
  2    +25.90ms *    +3.51ms  22.39ms  200  13K /test/ysf.png
sherif@fingolfin:~$ 

From above output, we can see all the resources that are starred are pushed by the server, both sets of resources pushed with 'Link' header and using the 'H2PushResource' are visible.
The resource ysf_096.png was pushed once, even though it was mentioned twice in the config, also resources pushed with early hint mechanism using 'H2PushResource' are sent using the HTTP 103 Early Hints response containing the Link header which could also be seen in the previous curl output.
Last test is to verify how browsers are handling server pushes.
This was done using Vivldi browser:


Using http/2 is gaining increasing popularity specially with CDN networks as it helps cutting down page load times by a big margin.
Dynamic applications using languages like php, Java and the like can always inject Link headers in their response to push large resources that might be still referneced deep inside the page or one of its dependancies to speed up page load time.



Thursday, 12 March 2020

Implementing a simple DNS service using dnsmasq

In many cases, one would need to have his own DNS service which would resolve local domain names into local IP address, and in the same time, still falls back to a public DNS like google DNS or OpenDNS name servers for external name lookups.

One easy way to implement this on Linux is to use the dnsmasq deamon to implement a local DNS server that fails over to a public one.

First, we need to install dnsmasq on the system and enable it, to do so we run the below commands:

[root@beren ~]#  yum install dnsmasq.x86_64 dnsmasq-utils.x86_64
[root@beren ~]# systemctl enable dnsmasq

[root@beren ~]# systemctl start dnsmasq


To check the if the service is running, we use systemctl status command and would expect to see an output similar to the one below:

[root@beren ~]# systemctl status dnsmasq
● dnsmasq.service - DNS caching server.
   Loaded: loaded (/usr/lib/systemd/system/dnsmasq.service; enabled; vendor preset: disabled)
   Active: active (running) since Thu 2020-03-12 18:21:35 EDT; 12s ago
 Main PID: 3253 (dnsmasq)
    Tasks: 1 (limit: 23977)
   Memory: 716.0K
   CGroup: /system.slice/dnsmasq.service
           └─3253 /usr/sbin/dnsmasq -k

Mar 12 18:21:35 beren systemd[1]: Started DNS caching server..
Mar 12 18:21:35 beren dnsmasq[3253]: started, version 2.79 cachesize 150
Mar 12 18:21:35 beren dnsmasq[3253]: compile time options: IPv6 GNU-getopt DBus no-i18n IDN2 DHCP DHCPv6 no-Lua TFTP no-conntrack ipset auth DNSSEC loop-detect inotify
Mar 12 18:21:35 beren dnsmasq[3253]: reading /etc/resolv.conf
Mar 12 18:21:35 beren dnsmasq[3253]: using nameserver 192.168.178.1#53
Mar 12 18:21:35 beren dnsmasq[3253]: read /etc/hosts - 2 addresses


To configure the now working name service, we need to edit the file /etc/dnsmasq.conf and add the the following configuration:

#Our config starts here, we listen to our local IP address:
listen-address=10.0.2.15

#Domain definition of the name server:
domain=sherif.local

#expand hosts file entries to be part of the domain:
expand-hosts
# Don't use resolv.conf for up stream servers
no-resolv

#Public DNS to fall back to.
server=8.8.8.8

address=/sherif.local/172.16.0.10
address=/beren.sherif.local/172.16.0.1
address=/dns.sherif.local/172.16.0.1
address=/mail.sherif.local/172.16.0.2
# Config end.


Then we need to restart the dnsmasq service for the config to take effect:

[root@beren ~]# systemctl restart dnsmasq

To test the new name service, we can use the nslookup command as below:

[root@beren log]# nslookup
> server 10.0.2.15
Default server: 10.0.2.15
Address: 10.0.2.15#53
> www.google.com
Server:         10.0.2.15
Address:        10.0.2.15#53

Non-authoritative answer:
Name:   www.google.com
Address: 172.217.17.68
Name:   www.google.com
Address: 2a00:1450:400e:80c::2004
> mail.sherif.local
Server:         10.0.2.15
Address:        10.0.2.15#53

Name:   mail.sherif.local
Address: 172.16.0.2
> dns.sherif.local
Server:         10.0.2.15
Address:        10.0.2.15#53

Name:   dns.sherif.local
Address: 172.16.0.1
>


As you can see, the mail server is able to resolve both local and external names successfully.
Last step now is to update our machine resolv.conf to point to the new DNS service which we run locally now.
To do this, we need to change the file /etc/resolv.conf to contain following line:

nameserver 10.0.2.15

Once done, we need to protect the file from Linux network manager service, one way to do this is to make the file immutable, we can use chattr to set the immutable attribute of the file as below:

[root@beren log]# chattr +i /etc/resolv.conf
[root@beren log]# lsattr /etc/resolv.conf
----i-------------- /etc/resolv.conf
[root@beren log]#



This way the networkmanager service will not be able to change our new DNS name server configuration upon system reboot.


Monday, 24 February 2020

Apache: How to conditionally inject Authorization header

 In one of the cases I was helping with, I ran into a requirement that needed to have a webhook be sent from Jira to Jenkins to trigger a build based on certain condition.
You can trigger a webhook from Jira based on a Jira search query or on a certain action, eg: upon creation of an issue in a project for example.

The problem is that Jira doesn't support using userinfo fields in the URL, which is actually deprecated and should be used as per the RFC https://tools.ietf.org/html/rfc3986, and thus, Jira would have a way to send authentication information out of the box to Jenkins to trigger a build.

To solve this, I proposed to run Jenkins behind an Apache reverse proxy and use a virtual host definition similar to the below:

<VirtualHost *:80>
     ServerName feanor
     DocumentRoot /var/www/html/

     SetEnvIf Remote_Addr "10.0.0.12" buildtriggerjira
     RequestHeader set Authorization "Basic c2hlcmlmOnNoZXJpZg=="   env=buildtriggerjira

     ProxyPreserveHost on
     ProxyVia on
     ProxyPass "/jenkins" "http://feanor:8080/jenkins"
     ProxyPassReverse "/jenkins" "http://feanor:8080/jenkins"
</VirtualHost>

The above configuration will inject an Authorization header in the incoming request on condition that the remote address is the IP address of the Jira server.
This works by setting an Apache environment variable to a certain value in the SetEnvIf condition and then setting the Authorization header if the variable has that value set.

More checks can be made to harden the condition by checking more request fields like the user-agent string or other headers set by the Jira http client sending the webhook request.

Sunday, 16 February 2020

Parallel execution in shell script using GNU parallel

Occasionally, one needs to look for certain strings inside very large files.
Very large file here is a text file that is more than 1GB in size.
The file could be opened by Linux tools such as less, but it becomes challenging to navigate the file and search inside it.
Also if we have a much bigger file, things can become more complex to handle with a single threaded application like less.
Things can become even more challenging if we have multiple files of that size, then one needs to either use grep and sed to look for patterns and do edits without the need to load the file(s) in memory.
Even with those command line filters, which mainly act on the input stream, doing a casual grep can take time to provide results.

To solve this, we would need to use the GNU parallel command to make use of CPU parallel execution.
parallel allows commands to be run in separate processes either on the same host or on multiple hosts, that can run in parallel processes, those would then do processing for parts of the input. parallel then collects the output of those processes and generates the result of the execution jobs.

A simple example for running parallel is below:

sherif@ulmo ~/logs $ < catalina.out parallel --pipe grep -C 4 'OutOfMemory'

The above command pipes the catalina.out that is redirected to standard input to the grep -C to look for the string.
The above form uses defaults block size (1MB) and the default number of jobs will be created (same as number of CPU cores available corresponding to 100%).

Using more optimization for the parallel command, we can achieve more efficiency compared to running a single thread filter, below is a comparison done for execution time for a standard egrep done on a 2GB file and same using parallel with 8 jobs and 100MB block size: 

[root@feanor ~]# time egrep -ni "^/ditto" all_files.txt |wc -l
82

real    0m5.401s
user    0m4.773s
sys     0m0.210s


[root@feanor ~]# time parallel --pipepart --joblog joblog1 -a all_files.txt --block 100M -P8 -k egrep -ni "^/ditto" |wc -l
Academic tradition requires you to cite works you base your article on.
When using programs that use GNU Parallel to process data for publication
please cite:

  O. Tange (2011): GNU Parallel - The Command-Line Power Tool,
  ;login: The USENIX Magazine, February 2011:42-47.

This helps funding further development; AND IT WON'T COST YOU A CENT.
If you pay 10000 EUR you should feel free to use GNU Parallel without citing.

To silence the citation notice: run 'parallel --bibtex'.

82

real    0m2.819s
user    0m7.224s
sys     0m1.801s
[root@feanor ~]#





As you can see, parallel was able to produce the same result in half the time taken by a single threaded egrep.
The fact that the user time is much higher for execution done by parallel, means that parallel made more use of CPU cores as the total CPU user time is higher than actual time (real) take by the process.
(for more information about the time command check this stackoverflow thread: https://unix.stackexchange.com/questions/40694/why-real-time-can-be-lower-than-user-time)

For more information and examples for using parallel, please check the excellent GNU parallel man page: https://www.gnu.org/software/parallel/man.html


Saturday, 8 February 2020

Usign Java Keytool to create a selfsigned SSL Certificate with SAN extentions

Recently I was doing some testing on Java Identity and Access management on premise tools and SAML integration with Apache mod-auth-mellon and other Java applications supporting SAML 2.0.
I needed to create a selfsigned SSL cert for a Java application so that I can ensure the whole SAML traffic is encrypted and secure.
Since I wanted to test multiple apps using multiple hostnames on the same server, I needed to have a generic certificate to save time.
Since I am also not a big fan of wildcard certs, and those also have the limitation of working only under a single top-level domain, I elected to use the Subject Alternative Name SSL certificate extension (SAN); so that I can define multiple definite none wildcard names and I can also have multiple domains and IP addresses if needed.

Creating a SAN SSL cert using keytool turned out to be straight forward, we just need to add the keytool option "-ext" and use the SAN name argument and then list all the SAN types needed.
In below example we use the DNS type and the IP type:

sherif@Luthien:~$ keytool -genkey -keyalg RSA -alias selfsigned -keystore keystore.jks -storepass password -validity 730 -keysize 2048 -ext san=dns:localhost,dns:localhost.com,dns:luthien,dns:redash.luthien.tst,dns:luthein.tst,ip:192.168.56.101

 The resulting certificate would look like below:

sherif@Luthien:~$ keytool -list -v -keystore keystore.jks
Enter keystore password: 
Keystore type: jks
Keystore provider: SUN

Your keystore contains 1 entry

Alias name: selfsigned
Creation date: Feb 8, 2020
Entry type: PrivateKeyEntry
Certificate chain length: 1
Certificate[1]:
Owner: CN=redash.luthien.tst, OU=Sherif Redash test, O=Sherif Tests, L=Cairo, ST=Cairo, C=EG
Issuer: CN=redash.luthien.tst, OU=Sherif Redash test, O=Sherif Tests, L=Cairo, ST=Cairo, C=EG
Serial number: 552e3728
Valid from: Sat Feb 08 14:48:35 CET 2020 until: Mon Feb 07 14:48:35 CET 2022
Certificate fingerprints:
     MD5:  C9:B6:DF:E7:3A:5E:EB:6C:97:C7:CD:8F:99:5A:9A:CD
     SHA1: EE:B7:C6:6F:C5:9C:15:36:5F:A7:12:95:4E:AA:8B:59:FB:08:B6:17
     SHA256: 25:3A:07:8D:FE:5B:FF:74:BE:E7:5F:EA:51:5E:8A:D8:31:41:7E:39:33:36:95:F0:27:C1:D8:2B:CE:54:FB:DA
Signature algorithm name: SHA256withRSA
Subject Public Key Algorithm: 2048-bit RSA key
Version: 3

Extensions:

#1: ObjectId: 2.5.29.17 Criticality=false
SubjectAlternativeName [
  DNSName: localhost
  DNSName: localhost.com
  DNSName: luthien
  DNSName: redash.luthien.tst
  DNSName: luthein.tst
  IPAddress: 192.168.56.101
]

#2: ObjectId: 2.5.29.14 Criticality=false
SubjectKeyIdentifier [
KeyIdentifier [
0000: F5 18 B2 15 5F 3F B9 63   E3 0F D3 8C 58 5A A0 67  ...._?.c....XZ.g
0010: C7 84 EF E1                                        ....
]
]

*******************************************
******************************************


As you can see, we have created the certificate with the SAN extensions that would allow us to use multiple names.

Please take a look at the Java keytool official documentation from Oracle: https://docs.oracle.com/javase/8/docs/technotes/tools/unix/keytool.html



Sunday, 27 October 2019

Regular experession matching in Postgres

Lately I ran across the need to extract user id text strings from a text colum in a database.
The text could be large and might contain multiple user id matches in the same row.

Thus I imported to the data into Postgres and started looking for Postgres functions that would help me use regular expressions to extract a text part of the column that matches the user id text.

I found a very useful Postgres function: REGEXP_MATCHES('input string', 'regexp pattern', 'flags').
This function works very similar to the sed Unix command, it applies the pattern against input string and allows the use of flag modifiers similar to the one used with sed.

Below is an example of a sample data:

The ids as you can see can be found in any part of the comment-text field.
A query using the  REGEXP_MATCHES function would help extract those Ids into a string arrary type defined in Postgres, the query looks like below:

select REGEXP_MATCHES("comment-text", '\[uidkey\:user[a-fA-F0-9]+\]','g') from comments where "comment-text" like '%[uidkey:%]%';

The output of the query looks like below:
As you can see the REGEXP_MATCHES function matched all 5 ids from the 2 rows that contain them and returned 5 array objects with 1 element each containing those ids.
To cast the array objects to standard text strings we use another Postgres function: ARRAY_TO_STRING(Array_object, 'array element separator', 'optional null element replacement'.

The final version of the query looks like below:

select array_to_string(REGEXP_MATCHES("comment-text", '\[uidkey\:user[a-fA-F0-9]+\]','g'),',')
from comments where "comment-text" like '%[uidkey:%]%';

The output of the query looks like below:
Thus, using those methods, we can extract text from text fields and build more complex queries based on that data.

Enjoy playing with your data !!






Tuesday, 19 March 2019

Setting RequestHeader in Apache

Apache can be used to inject a Request Header in the incoming request that can be either consumed by Apache or forwarded further to another underlying service, in this case Apache works essentially as a reverse proxy.

In a test setup where Apache works as a reverse proxy in front of tomcat, the below Apache configuration is used to implement the reverse proxy functionality and add a Request Header:

<VirtualHost *>
   <Location "/sherif">
      ProxyPass http://127.0.0.1:8080/sherif
      ProxyPassReverse http://127.0.0.1:8080/sherif
      RequestHeader set myh "valueofarequestheader"
   </Location>
</VirtualHost>

In this setup, tomcat is using default port 8080, it has a defined context path for a dummy application defined in tomcat under its context.xml:

<Context>
    <WatchedResource>WEB-INF/web.xml</WatchedResource>
    <WatchedResource>WEB-INF/tomcat-web.xml</WatchedResource>
    <WatchedResource>${catalina.base}/conf/web.xml</WatchedResource>

    <Manager pathname="/sherif" />
</Context>


Under tomcat webapps, we have folder created as sherif and has a simple index.html file to support the test:
[root@localhost conf]# ls -lt ../webapps/sherif/
total 4
-rw-r--r-- 1 root root 14 Mar 19 16:27 index.html
[root@localhost conf]#

To verify the header being added, we configure tomcat to log the head myh.
This is done on tomcat server.xml accesslog value as below:

<Valve className="org.apache.catalina.valves.AccessLogValve" directory="logs"
               prefix="localhost_access_log" suffix=".txt"
               pattern="%h %l %u %t &quot;%r&quot; %s %b %{myh}i" />

Once a request is sent to http://localhost/sherif, tomcat logs the below log showing the request header being added by Apache and reaching tomcat:

127.0.0.1 - - [19/Mar/2019:16:46:47 -0400] "GET /sherif/ HTTP/1.1" 200 17 valueofarequestheader

This configuration is useful in passing headers to backend services in case those are not already sent by the source user agent.