Apr 13, 2014

Akka Java for large-scale event processing

We are designing a large scale distributed event-driven system for real-time data replication across transactional databases. The data(messages) from the source system undergoes a series of  transformations and routing-logic before reaching its destination. These transformations are multi-process and multi-threaded operations, comprising of smaller stateless steps and tasks that can be performed concurrently. There is no shared state across processes instead, the state transformations are persisted in the database, and each process pulls its work-queue directly from the database. 

Based on this, we needed a technology that supported distributed event processing, routing and concurrency on the  Java + Spring platform, the three options considered were, MessageBroker (RabbitMQ), Spring Integration and Akka

RabitMQ: MQ was the first choice because it is the traditional and proven solution for messaging/event-processing. RabbitMQ, because it is popular light-weight open source option with commercial support from a vendor we already use. I was pretty  impressed with RabbitMQ, it was easy to use, lean, yet supported advance distribution and messaging features. The only thing that it lacked for us, was the ability to persist messages in Oracle.  

Even though RabbitMQ is Open Source (free), for enterprise use, there is a substantial cost factor to it. As MQ is an additional component in the middleware stack, it requires dedicated staff for administration and maintenance, and  a commercial support for the product. Also, setup and configuration of MesageBroker has its own complexity and involves cross-team coordination.

MQs are primarily EAI products and provide cross-platform (multi-language, multi-protocol) support. They might be too bulky and expensive when used just as asynchronous concurrency and parallelism solution.

Spring Integration:  Spring has a few modules that provide scalable asynchronous execution.
Spring TaskExecutor  provides asynchronous processing with lightweight thread pool options.
Spring Batch  allows distributed asynchronous processing via the Job Launcher and Job Repository. 
Spring Integration extends it further by providing EAI features, messaging, routing and mediation capabilities.

While all three Spring modules have some of the required feature, it was difficult to get everything together. Like this user, I was expecting Spring Integration would have RMI-like remoting capability.

Akka Java: Akka is a toolkit and runtime for building highly concurrent, distributed, and fault tolerant event-driven applications on the JVM. It has a Java API and I decided to give it a try.

Akka  was easy to get started, I found Activator quite helpful. Akka is based on Actor Model, which is  a message-passing paradigm of achieving concurrency without shared-objects and blocking. In Akka, rather than invoking an object directly, a message is constructed and send it to the object (called an actor) by way of an actor reference. This design greatly simplifies  concurrency management. 

However, the simplicity does not mean that a traditional lock-based concurrent program (thread/synchronization) can be  converted into Akka with few code changes. One needs to design their Actor System by defining smaller tasks, messages and communication between the them.  There is a learning curve for Akka’s concepts and Actor Model paradigm. It is comparatively small, given the complexity of concurrency and parallelism that it abstracts.

Akka offers the right level of abstraction, where you do not have to worry about thread and synchronization of shared-state, yet you get full flexibility and control to write your custom concurrency solution.   

Besides  simplicity, I thought the real power of Akka is, remoting and its ability to  distribute actors across multiple nodes for high scalability. Akka's Location Transparency and Fault Tolerance make it easy to scale and distribute  application without code changes. 

I was able to build a PoC for my multi-process and multi-threading use-case, fairly easily.  I still need to work out Spring injection in Actors.

A few words of caution, Akka’s Java code has a lot of typecasting due to Scala’s type system and achieving object mutability could be tricky. I am tempted to reuse my existing JPA entities (mutable) as messages for reduced database calls.
Also, Akka community, is geared towards Scala and there is less material on Akka Java.

In spite of all this, Akka Java seems cheaper, faster and efficient option out of the three.

Feb 13, 2014

Setup RabbitMQ and Pika Python client on MacOS

There is an excellent tutorial on RabbitMQ,  however, I thought it lacked detailed steps on installation and setup of RabbitMQ server and Pika Python client. I would like to share the steps on MacOS.

Installing and Running RabbitMQ


RabbitMQ is available in homebrew

brew install rabbitmq

once installed, run the server

sudo  /usr/local/sbin/rabbitmq-server

verify RabbitMQ is running 

http://localhost:15672/ should give a login prompt, login using guest/guest


Installing Pika

Install python if you don't already have it,

brew install python

The Python formula comes with Pip and setuptools, check

/Library/Python/<yourversion>/site-packages

install pika

sudo pip install pika=0.9.8

I also ran

sudo easy_install pika

it pulls-in additional packages for pika.

That is it, you should be able to use Pika in your python programs.

You should be all set now and can start with the RabbitMQ tutorial.


Jun 30, 2013

Mocking REST services using Apache HTTPD

We needed the ability to run our application using mock-data, but due to the complexity of our data-sources, we could not mock data or data sources. We did however have RESTful web services that exposed the domain data as Web Resources, these services are invoked by the Application Tier to present UI. We found that it was easier to mock our REST services than having to create a new data source with mock data.

We configured Apache Webserver ( HTTPD) to simulate the REST resources with the resource URLs that matches the actual service URLs. For a service resource: http://servicehost:8080/services/mydomain/users/test01/profile, that returns a JSON object for the user's profile, we created a "profile" file, containing the desired JSON, under /htdocs/services/mydomain/users/test01. The file can be accessed via the URL: http://httpdhost:80/service/mydomain/users/test01/profile that matches the actual service URL except for hostname and port.

Apache HTTPD was configured to return response content-type as application/json instead of default text/plain for the required directories. This was done by setting the httpd conf ForceType for a Location to set a particular content type.
<Location /services>
  ForceType application/json
</Location>

So, by changing the external parameter for the services URL (hostname and port) to a mock service provider (HTTPD), one can run the application using mock data.

We found this simple approach particularly useful in our development environment where there is an independent UI development team. UI developers can now develop and test the UI using desired mock data without being dependent on services to be available or having to create a mock data at each service-call level. 
It's easier for them to manage and change JSON (JavaScript Object Notation) as oppose to mock-data approach where they would need to understand the underlying data source and data representation. Also, the mock configuration is managed as single switch ie an  externalized application parameter (service URL), as oppose to changing all the data-stores (service URLs) in the UI code.

I got some good insight and validation for the REST service design when I started creating the resources in htdocs directory structure.

May 15, 2013

Remote Client Web Performance Monitor

Selenium is an excellent open source tool for automated functional testing of web applications. We use it along with JBehave for automation and  Behavior Driven Development of our large-scale browser-based application.  The tests are run from Atlassian Bamboo and we use Sauce Labs for the cross browser testing. This setup allows for a full regression testing of the application on all supported browsers with very little manual testing. This automation, along with high-coverage unit-test suite (Junit/Mockito) and puppet-based one-click deployment, is core to our continues improvement & delivery platform.

We are also using selenium to measure Page Load time of our Single Page Web application (ExtJS4.0).  Selenium WebDriver provides more reliable page load time then Selenium RC, as it uses browser's native support for automation as oppose to injecting JavaScript.
This extends Selenium ability to act as an automated performance testing tool, it can  validate ( or just log) performance of the application, along with performing functional validation.

We went a step further with Selenium implementation, were we used its WebDriver to run tests on any remote client browser to understand the real performance of the application. The user is given a test URL that runs the application via Selenium webdriver on the client browser and collect required diagnostic data.

Selenium has a remote server, which allows the selenium tests to run on remote browser. However, the remote server needs to be installed and running on the client machine, and we did not want that. We needed a reverse solution, where selenium tests run on a remote server and the application runs on the local browser.

This was solved by creating a simple web app (servlet) that can identify  its browser (using user-agent request header) and based on browser type, invoke appropriate webDriver. It would then run the application and collect load time metrics. This metric is sent via HTTP/JSON to CouchDB. I am a big fan of  CouchDB, it provides rapid application development and  schema flexibility for tool’s feature extension. 

I am searching a way to reuse my existing Selenium scripts for Load Testing, maybe run them through JMeter and capture response time in CouchDB. For now, we are using HP LadRunner with AJAX extension and it is functioning well.


Apr 30, 2013

Client-Tier Performance Tuning of Single Page Applications

JavaScript (Browser-based/Single Page) applications are new generation applications that tap on the browser’s processing power to provide rich user experience and fast user interaction. However, these applications do have slower initial page load then traditional applications. As  client-tier is handling more processing in these apps, it plays a bigger role in application responsiveness; client-factors like browser, network, processing power can slow down a request even before it ever reaches the server, even a well-performing back-end, might give a bad performance to certain clients. Along with server optimization ( lean architecture), these applications require client-tier performance tuning to have fast application response time. 


A browser’s main function is to get requested web resources, parse the received content, execute JavaScript and display the content in its window. It has a browser engine, a rendering engine, a JavaScript engine and networking component to fetch content from server resources via HTTP requests. The HTTP request to fetch a resource involves a DNS name lookup, SSL negotiations, setting up TCP connections, transmitting HTTP requests, content download, and fetching resources from cache. Browser processing (rendering/JavaScript execution) is a single threaded process, except network operation, which is done in parallel. However, browsers do have a limit on the number of parallel network connections to a domain, varies from 2 -6 connections. Browsers have their own implementation of these components which is why response time and rendering of same content differs based on browsers.

In order for JavaScript to be executed on the browser it first needs to be transferred JavaScript source code from the server to the browser. This not only cause delays due to network latency, but also leads to synchronous execution of the page. 

When browser processes a downloaded JavaScript resource (script tag), it blocks all other JavaScript/CSS resources till that particular JavaScript is parsed and executed. 
This becomes a major issue when using AJAX toolkit that have large JavaScript libraries.
Also, browser-based applications have large application specific JavaScript code. Usually this code is modularized into smaller more manageable JS files. This means there is a large JavaScript code, distributed in small JS files that need to be transferred from the server to the browser.

So download and processing of JavaScript code itself increases network latency and can be a major bottleneck during page load.

The most efficient solution for this is to combine multiple JavaScript/CSS files into a single large file, which is compressed and minified before being transferred to the browser. 

There are several frameworks available for this, we used JAWR. JAWR provides server side configuration to combine multiple files (JS or CSS) into a single file called bundle. These bundles are generated during server startup and are invoked by calling JAWR tag libs replacing <script> tags. JAWR also applies minification and compression to the bundles. 

We bundled all of the application JS code together with ExtJS into a single JS bundle. 
We created another bundle for our client-side translation code: I8N JavaScript resource bundles and ExtJS locales. The bundle for each language had to be defined separately, otherwise properties would be overwritten. JAWR does have an I8N message generator that can optimize translation bundling, but we did not use it because we had a custom solution for client-side translation. 
We used OWA for client-side event tracking. It had few JavaScript files that need to be included in the application pages to allow OWA to track the events on that page and send the events to a centralized OWA server. We had to do some custom coding, but were able to include OWA JS in our application JAWR bundles. 
We had a single CSS bundle for all the CSS files in the application. This reduced the number of sever calls from the page and also there was less blocking for JS execution. 
We saw a major performance boost to application page load with this change.

The next performance tuning was to manage the number of HTTP requests from the page. As browsers have limit to maximum parallel requests that it can send to a single domain, this can block the AJAX requests on the page and cause delays.

Typically, most of resources on a web page are images, so we need to find ways to reduce network calls for images. The easiest approach is to cache images, because they rarely change. We configured cache control header to cache images and other non-changing static content on the browser. We also use Content Delivery Network –Akamai as a Web Proxy cache and for Geo-accelerates content delivery.

For no caching (first application run) situation, one of the approaches  to reduce the number of images related HTTP calls that we tried was JAWR Image Sprite, we faced few UI complexities with this approach and did not use it.

We were severing static content and dynamic content from the same domain, so we tried to move the images into a separate domain (cookie-less and non-SSL). This lead to mixed content (HTTPS/HTTP) issue, which is not secure and certain browser presents a user prompt for this. Changing image's domain to be secure (HTTPS) increased the SSL negotiation time which negated any gain that we got by increasing parallel network calls from the browser.

We finally implemented a solution to  perloaded/ pre-cache the images in a previous page (login page). The images are preloaded and avaliable in the browser cache before the page that displayed them was called. This reduced the number of HTTP calls on the main page. Also we moved the images to an un-authenticated host, which reduced load on the back-end authenticated server and we reduced a network hop. This gave us a major performance improvement, especially in IE browsers. This might not a generic solution, but I guess using Image Sprites should give the same results.

The images need to be optimized during UX design time, to have smaller size and proper format. However, at times, one might get more latency for smaller images than for larger ones. This can happen when the client gets asymmetric bandwidth and the content response size is not proportionately bigger than the request size as expected by asymmetric (upload: download) ratio.

Along with these changes, we had also tuned back-end (HTTPD/tcServer) to handle large number of concurrent HTTP requests by configuring KeepAlive, Connection timeout and maxThreads/maxClients. On tcServer, we configure Non-Blocking IO connector which is optimized to handle large number of concurrent HTTP connections.

With all these optimizations, our large scale enterprise web application now has a high-performing and rich front-end with lean and scalable back-end.

Note: Checkout Google Chrome Frame if you need to provide new functionlity on legacy browsers.

Mar 13, 2013

Custom MBean configuration for Connection Pool monitoring

We ran into some interesting problems in configuring Hyperic to monitor JDBC connection pool via JMX. I would like to share the details, hoping it would save time for others.

When the connection pool and data sources are configured in the container using JNDI, Hyperic is able to identify pool MBeans automatically. However, if the pool configuration is within the application (war) and not in the container, MBeans will not be visible in Hyperic. To make them available in Hyperic one will need to configure JMX plugin for Hyperic.

Depending on the connection pool, you might need to register the MBeans explicitly. Some connection pools (c3p0) have auto-registered MBeans, so if you use them, the MBean will be auto-registered and become available to JMX tools like JConsole. However, these MBeans have dynamic object name that changes on server restart. This would be a problem in cases where you need a static MBean name to be able to identify it for setting up monitors and alerts. Hyperic does have a patch to allow monitoring MBeans with dynamic name http://communities.vmware.com/thread/389027, but this patch needs to be installed on the enterprise tcServer installation. 
If you do not wish to use the patch, you will need to register connection pool MBean explicitly with a static name.

Registering (export) the MBean with JMX server can be tricky and can lead to sever connection leak problem. One needs to be careful in selecting the MBean that should be exposed. Typically, the datasource beans itself has all the required connection attributes that one needs for monitoring like number of connection Idle/Used. But this datasource bean also has additional methods like getConnection*, that actually borrows a connection from the pool. 
If this data source is directly registered, and then monitored by JConsole (or likes), it may cause these getConnection* methods to be invoked and thus lead to connections being  borrowed from the pool without ever being released. Tools like JConsole are configured to display all the exposed attributes for the registered MBean in its attribute tab by calling the assessors  to get the attribute value. If the bean had assessors like getConnection*, Jconsole will invoke them, causing connections to be borrowed.

To avoid this connection leakage problem and have a static MBean name that can be used for monitoring and alerts, create a custom MBean. Configure this custom bean to have just the attributes that are needed for monitoring and alerts (caution: do not extend the data source bean).

Here is the JMX configuration for a tomcat - pool that is working well for us through JConsole and Hyperic.

       <beans:bean id="exporterOracle"   class="org.springframework.jmx.export.MBeanExporter" lazy-init="false">
              <beans:property name="beans">
                     <beans:map>
                           <beans:entry key="bean:name=OracleConnectionMBean" value="#{dataSource.getPool().getJmxPool()}"/>        
                     </beans:map>
              </beans:property>
       </beans:bean>

<beans:bean id="dataSource" class="org.apache.tomcat.jdbc.pool.DataSource "
              destroy-method="close">
              <beans:property name="driverClassName"   value="${oracle.db.driver}" />
              <beans:property name="url" value="${oracle.db.url}" />
              <beans:property name="username" value="${oracle.db.user}" />
              <beans:property name="password" value="${oracle.db.password}" />
              <beans:property name="jmxEnabled" value="true" />
              <beans:property name="initialSize" value="${oracle.pool.initialsize}" />
              <beans:property name="maxActive" value="${oracle.pool.maxactive}" />      
              <beans:property name="maxIdle" value="${oracle.pool.maxIdle}" /> <!-- maxIdle needs to be same as maxActive -->
              <beans:property name="testOnBorrow" value="true"></beans:property> <!-- Needs to be set to prevent Connection timeout exception -->         
              <beans:property name="validationQuery" value="select 1 from dual" /> <!-- This needs to be set if testofBorrow is set -->
       </beans:bean>

Check that the MBean name is static "bean:name=OracleConnectionMBean" and bean that is exposed is specific JMXPool interface of the datasource (dataSource.getPool().getJmxPool()), not the datasource itself.


Feb 15, 2013

Protecting Web Resources


We are working on a modernization program, as part of which we have built a new enterprise Portal that replaces existing Websphere Portal. This new Portal is a simple Web20 application with ExtJS4.0/Spring/Restful services running on the VMware Tc Server.

This migration required replacing Webspere Portal components with light-weight open source products.  However, WebSEAL was one product that was impressively lightweight and helped us solve some complex integration problems.

WebSEAL is the resource manager that acts as a reverse Web proxy. It receives HTTP/HTTPS requests from a Web browser and delivering content from its own Web server or from junctioned back-end Web application servers. Web Requests passing through WebSEAL are authorized by the Tivoli Access Manager.


We were already using WebSEAL for authentication, single sign-on and high level HTTP URL authorization. In new Portal, we extended its use for integrating third party application with Portal.

Websphere Portal (WSP)  used WebClipper technology  to integrate and render third-party applications. WebClipper runs from within the portal server and manage session and identity across Portal and third-party apps. This solution was highly complex and created tight-coupling between app and the Portal.
In the new Portal we replaced WebClipper with WebSEAL-based integration, where we created a WebSEAL junction for the application, and used the secured junction URL to be rendered the app as  ExtJS Tabs within the Portal. WebSEAL provided secure rendering and session management for the app & Portal. WebSEAL also passed authentication token to the app, as trust HTTP Headers, eliminating the need to engage Portal server. 

We even used WebSEAL as an operational tool to control user traffic at
 run time.  We plan to use WebSEAL to display error pages or redirect traffic to a specific web resources. While this can be done on any webserver, WebSEAL allows the redirection based on fine grain permissions/ACLs.

WebSEAL is an excellent product, I am not sure if there is any open source product that can provide the same capabilities. Its only limitation is that it is tightly coupled with Tivoli Access management and enforces security policies against just Tivoli Access Manager, also that it is not open source.

Spring Security does provide similar web security against any policy manager and can be configured as an HTTP reserve proxy. However, Spring Security protects web resources within its own web context and needs an application server. If you need to protect third-party web resources/application using Spring security, you will need to stream the HTTP request and attach security headers to it. For protecting simple webservices  this approach will work. But for protecting web applications, streaming can create complexities. Moreover, you will need to serve application’s static content through the app server.

As the trend of HTTP interface grows, products  provide HTTP/JSON interfaces, ex Solr, CoucbBase , so will the need to non-instrusively secure web-resources . I would like to know how others are solving this problem.


Jan 9, 2013

Sonar with mysql


I wanted to share few details on configuring Sonar to use mysql. It is quite easy, just follow documentation available at  http://docs.codehaus.org/display/SONAR/Installing+Sonar

On sonar startup, if you get this error 
Cause: org.apache.commons.dbcp.SQLNestedException: Cannot create PoolableConnectionFactory (Unknown database 'sonar')

It might mean that sonar database is not configured properly. Try creating the database by running the scripts  available @ https://github.com/SonarSource/sonar/tree/master/sonar-application/src/main/assembly/extras/database/mysql.

Now on maven side , you need to add this to the setting.xml
<profile>
            <id>sonar</id>
            <activation>
                <activeByDefault>true</activeByDefault>
            </activation>
            <properties>
                <!-- SERVER ON A REMOTE HOST -->
               <sonar.jdbc.url>jdbc:mysql://localhost:3306/sonar?useUnicode=true&amp;characterEncoding=utf8</sonar.jdbc.url>
                <sonar.jdbc.driverClassName>com.mysql.jdbc.Driver</sonar.jdbc.driverClassName>
                <sonar.jdbc.username>sonar</sonar.jdbc.username>
                <sonar.jdbc.password>sonar</sonar.jdbc.password>
                <sonar.host.url>http://localhost:9000</sonar.host.url>
            </properties>
 </profile>


 If you get the error ‘Cannot load JDBC driver class’ on running mvn sonar:sonar  
 check for the driver class name. 
If driver class is derby or H2, it means the mysql configuration is not picked by from setting.xml. Try giving profile for sonar explicitly and including the path to setting.xml, this would ensure that the mysql settings is pulled up by maven.

If however the driver class name in the error is 'com.mysql.jdbc.Driver', this means that the profile is read properly. This error is mostly because of incorrect JDBC url; eventhough the eror messages makes you think that you need to copy the driver file somewhere for maven to pickup. 

In my case I had my jdbc url as jdbc:mysql://localhost:3306/sonar, when I changed to jdbc:mysql://localhost:3306/sonar?useUnicode=true&amp;characterEncoding=utf8 everything worked fine.