United States
20K followers 500+ connections

Join to view profile

About

I'm a Senior Staff Software Engineer with 12+ years of experience working on products…

Articles by Sandish Kumar

Activity

20K followers

See all activities

Experience & Education

  • Hewlett Packard Enterprise

View Sandish Kumar’s full experience

See their title, tenure and more.

or

By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.

Licenses & Certifications

Publications

Projects

  • Yardstick­ Spark

    Technologies Used:​ Spark ​ , Spark SQL ​Core​Data Frame's​, Spark ​Streaming​, Scala, HDFS, Apache Ignite, Yardstick Tool, D3js, ​Tableau​, AWS
    Data Sets: ​10 million twitter data records and 1 billion auto generated records.
    Description:
    Yardstick Apache Spark is a set of Apache Spark benchmarks written on top of Yardstick
    framework.

    I have successfully written Spark ​CoreRDD​ application to read auto generated 1 ​

    Technologies Used:​ Spark ​ , Spark SQL ​Core​Data Frame's​, Spark ​Streaming​, Scala, HDFS, Apache Ignite, Yardstick Tool, D3js, ​Tableau​, AWS
    Data Sets: ​10 million twitter data records and 1 billion auto generated records.
    Description:
    Yardstick Apache Spark is a set of Apache Spark benchmarks written on top of Yardstick
    framework.

    I have successfully written Spark ​CoreRDD​ application to read auto generated 1 ​
    billion​ records and compare with ​IgniteRDD​ in Yardstick framework to measure
    performance of Apache Ignite RDD and Apache Spark RDD.

    I have successfully written Spark ​DataFrame​ application to read from ​HDFS​ and analyze
    10 million twitter records using Yardstick framework to measure performance of Apache
    Ignite SQL and Apache Spark DataFrame.

    I have successfully written Spark ​Streaming​ application to read streaming twitter data and
    analyze twitter records in ​real ­time​ using Yardstick framework to measure performance of
    Apache Ignite Streaming and Apache Spark Streaming.

    Implemented ​test cases​ for Spark and Ignite functions using scala as language.

    Hands-­on experience in setting up 10 node Spark cluster on Amazon Web Service’s ​
    using Spark EC2 script.

    Implemented ​ and ​D3.js​ Tableau​ charts to show performance difference between Apache
    Ignite and Apache Spark.

    Other creators
    See project
  • Comparative Analysis of Big Data Analytical Tools –(Hive,Hive on Tez,Impala,SparkQL,Apache Drill, BigQuery, PrestoDB running on the Google Cloud and AWS)

    Technologies Used: Hive,Hive on Tez,Impala,SparkSQL, Apache Drill, BigQuery, PrestoDB, Hadoop, Cloudera
    CDH, Hortonworks HDP, Google Cloud Platform, Amazon web service
    Data set used: Twitter streaming data
    a. Installation of Hive,Hive on Tez,Impala, SparkQL, Apache Drill, BigQuery, PrestoDB, Hadoop,
    Cloudera CDH, Hortonworks HDP
    b. Schema design for data sets on all Hive,Hive on Tez,Impala, SparkQL, Apache Drill, BigQuery,
    PrestoDB
    c. Query design for given data set
    d…

    Technologies Used: Hive,Hive on Tez,Impala,SparkSQL, Apache Drill, BigQuery, PrestoDB, Hadoop, Cloudera
    CDH, Hortonworks HDP, Google Cloud Platform, Amazon web service
    Data set used: Twitter streaming data
    a. Installation of Hive,Hive on Tez,Impala, SparkQL, Apache Drill, BigQuery, PrestoDB, Hadoop,
    Cloudera CDH, Hortonworks HDP
    b. Schema design for data sets on all Hive,Hive on Tez,Impala, SparkQL, Apache Drill, BigQuery,
    PrestoDB
    c. Query design for given data set
    d. Debugging on Hive,Hive on Tez,Impala, SparkQL, Apache Drill, BigQuery, PrestoDB, Hadoop,
    Cloudera CDH, Hortonworks HDP
    e. Time Comparison of each Hive,Hive on Tez,Impala, SparkQL, Apache Drill, BigQuery, PrestoDB
    f. Time comparison between different cloud platforms
    g. Times metrics web based visualization design on google charts

    Other creators
    See project
  • Truck Events Analysis

    Technologies Used: Hadoop, HDFS, Hive, HBase, Kafka, Storm, RabbitMQ WebStormp, Google Maps
    Data Sets Used: - New York City Truck Routes from NYC DOT. -Truck Events Data generated using a custom
    simulator. - Weather Data, collected using APIs from Forcast.io. -Traffic Data, collected using APIs from
    MapQuest.
    a. Written a simulator to send/emit events based on NYC DOT data file.
    b. Written Kafka Producer to accept/send events to Kafka Producer which is on Storm Spout
    c…

    Technologies Used: Hadoop, HDFS, Hive, HBase, Kafka, Storm, RabbitMQ WebStormp, Google Maps
    Data Sets Used: - New York City Truck Routes from NYC DOT. -Truck Events Data generated using a custom
    simulator. - Weather Data, collected using APIs from Forcast.io. -Traffic Data, collected using APIs from
    MapQuest.
    a. Written a simulator to send/emit events based on NYC DOT data file.
    b. Written Kafka Producer to accept/send events to Kafka Producer which is on Storm Spout
    c. Written Storm topology to accept events from Kafka Producer and Process Events
    d. Written Storm Bolt to Emit data into Hbase, HDFS, RabbitMQ WebStomp
    e. Hive Queries to Map Truck Events Data , Weather Data, Traffic Data

    Other creators
  • Cimbal/MobApp Pay

    Cimbal is a mobile promotion and payment network designed to increase business sales and deliver targeted deals to
    consumers.
    Technologies used: Hadoop MapReduce, Hbase , Spring Data Rest Web Service, CDH
    Data Set : Users Payment Data
    a. Written MapReduce programs to validate the data
    b. Written more than 50 Spring Data Hbase rest API's in Java
    c. Schema design on Hbase and cleaning data
    d. Written Hive queries for analytic's on users data.

    Other creators
    See project
  • Social Media Sentiment Analysis Update on each 4 hours (Positive Negative Review's)

    Social Media Sentiment Analysis project gives positive negative reviews on our favorite words search, it is based on Twitter tweets, Facebook status's etc... it has been implemented using Social media API's with BigData Technologies like Flume to get social media stream data regularly into HDFS and Cassandra, MongoDB to store sentiment analyzed data. and in this project i am using apache oozie to develop work flow and it will be ruing on each 4 hours. this is project is done with my self…

    Social Media Sentiment Analysis project gives positive negative reviews on our favorite words search, it is based on Twitter tweets, Facebook status's etc... it has been implemented using Social media API's with BigData Technologies like Flume to get social media stream data regularly into HDFS and Cassandra, MongoDB to store sentiment analyzed data. and in this project i am using apache oozie to develop work flow and it will be ruing on each 4 hours. this is project is done with my self interest.

  • Whole Genome Sequencing - Quality Check (QC)

    -Entire program was developed in Hadoop MapReduce praradigm
    -Performed custom Quality Check on genomic data to retain good quality reads
    -Implemented novel features in the software:
    -capability to handle sequencing-machine errors
    -automatic detection of base-line PHRED score
    -being platform agnostic and able to handle various file-formats i.e. from Illumina, 454 Roche, Complete Genomics, ABI Solid input format data

    See project
  • Whole Genome Sequencing - Sequence Alignment

    - Developed a Hadoop MapReduce program to perform sequence alignment on NGS data.
    - The MapReduce program implements algorithms such as Borrows-Wheeler Transform (BWT), Ferragina-Manzini Index (FMI), Smith-Waterman dynamic programming algorithm using Hadoop distributed cache.

    See project
  • Seismic Data Server & Repository (SDSR™)

    Our Seismic Data Server & Repository solves the problem of delivering, on demand, precisely cropped SEG-Y files for instant loading at geophysical interpretation workstations anywhere in the network. Based on Hadoop file storage, Hbase™ and MapReduce technology, the Seismic Data Server brings fault-tolerant petabyte-scale store capability to the industry. Seismic Data Server supports post-stack traces now with pre-stack support to be released shortly.

    Other creators
    See project
  • DDSR (Drilling Data Search and Repository):

    This project aims to provide analytics for Oil and Gas exploration data. This DDSR repository builded by using HBase, Hadoop and its sub projects. We are collecting thousands of wells data from across the globe. This data is stored in Hbase and Hive by using Hadoop MapReduce jobs. On top of this data we are building analytics for search and advanced search.

    Other creators
    See project
  • E Commerce (Obesessory.com)

    -

    Tech: Hsearch(Hbase+lucene), Hbase, Cassandra, Hive, Spark, Hadoop, MapReduce,Amazon Web
    Service, Linode, CDN, Scala, Java
    Data: Affilats feeds Rakuten, CJ, Affiliate window, WebGains.

    Pre-Processing:
    Crawling of 100+ sites Data using Nutch
    Fashion based ontology maintenance
    Using Spark, Scala:
    to Enriched given data using Fashion Ontology
    to Validation/Normalizing the data
    Designed schema, Modeling the data and Written program to store all validated data in…

    Tech: Hsearch(Hbase+lucene), Hbase, Cassandra, Hive, Spark, Hadoop, MapReduce,Amazon Web
    Service, Linode, CDN, Scala, Java
    Data: Affilats feeds Rakuten, CJ, Affiliate window, WebGains.

    Pre-Processing:
    Crawling of 100+ sites Data using Nutch
    Fashion based ontology maintenance
    Using Spark, Scala:
    to Enriched given data using Fashion Ontology
    to Validation/Normalizing the data
    Designed schema, Modeling the data and Written program to store all validated data in Cassandra
    Spring Data Cassandra Programs for Validation/Normalizing/Enriching and REST API to Develop UI Based manual QA Validation.
    Used SPARQL, Scala to running QA based SQL queries.

    Indexing:
    MR Programs on Hbase:
    To Standardize the Input Merchants data
    To Upload images to RackSpace CDN
    To Index the given Data sets into Hsearch
    To MR Programs on Hbase to extract the color information from Images including density
    To MR Programs on Hbase to Persist the Data on Hbase tables
    above MR jobs will run based on timing and bucketing

    Color-Obsessed:
    Using Image color and density data
    User will be allowed to select 1,2.. colors with different densities and result will be a list of
    products where each product image contains all give colors with exact density
    Written Hbase Spring rest web service for Color Obsessed Search API

    Post-Processing:
    Setting up the Spark Streaming and Kafka Cluster
    Developed a Spark Streaming_Kafka App to Process Hadoop Jobs Logs
    Kafka Producer to send all slaves logs to Spark Streaming App
    Spark Streaming App to Process the Logs with given rules and produce the Bad Images,Bad
    records, Missed Records etc
    Spark Streaming App collect user actions data from front end
    Kafka Producer based Rest API to collect user events and send to Spark Streaming App
    Hive Queries to Generate Stock Alerts, Price Alerts, Popular Products Alerts,
    New Arrivals for each user based on given likes, favorite,shares counts information
    Working on SparkMLlib for Recommendations, Coupons Recommendations, Rules
    Engine

    See project

Languages

  • English

    Professional working proficiency

  • Telugu

    Native or bilingual proficiency

  • Kannada

    Native or bilingual proficiency

  • Tamil

    Limited working proficiency

Recommendations received

4 people have recommended Sandish Kumar

Join now to view

View Sandish Kumar’s full profile

  • See who you know in common
  • Get introduced
  • Contact Sandish Kumar directly
Join to view full profile

Other similar profiles

Explore top content on LinkedIn

Find curated posts and insights for relevant topics all in one place.

View top content

Add new skills with these courses