Python developer to use multithreading and multiprocessing to maximise Linux server capacity

Job ID: 34910190

Budget: €30 – €250 EUR

We’ve developed a python pipeline for calculating counts in next generation sequencing libraries that have been restriction digested in two (one sequence split in two). This requires the calculation of the outer product from a pair of count files. Outer product matrices (~34Gb) are stored in partitions as NumPy files and are read in to calculate averages and standard deviations for sets of libraries from two conditions. These are then used to find the outer product entries with the highest z-score between the two conditions.

Currently the pipeline is general use and accepts a set of parameters at the start and a sample table that specifies the name of library files to be read in and which condition they are from. The pipeline in total takes around 5 hours but is only running on one core. We’ve briefly tried to use multiprocessing libraries available to utilize multiple cores, but this doesn’t speed up the pipeline.

We are running this on a Debian server with 126Gb RAM and 32 cores. We want to use the full server capacity. We’ve a test set of data that runs quicker if needed to run simulations.

The ideal candidate is an advanced python user with an in-depth knowledge of multiprocessing python through a Linux server with the multiprocessing library, but not with Dask for now. You will be dealing mainly with our bioinformatician based in France, but also with the server managed based in Canada.

We’ll potentially have more freelance work available in the future with some projects involving Python, Perl and pyCUDA. Candidates with this skill set will be preferential.
Related categories: Perl Python Linux CUDA