About
Articles by Sandish Kumar
Activity
20K followers
Experience & Education
Licenses & Certifications
Publications
-
Visualizing NetFlow Data with Apache Kudu, Apache Impala (incubating), StreamSets Data Collector, and D3.js
phdata
See publicationVisualizing NetFlow Data with Apache Kudu, Apache Impala (incubating), StreamSets Data Collector, and D3.js
-
Visualizing NetFlow Data with Apache Kudu, Apache Impala (incubating), StreamSets Data Collector, and D3.js
StreamSets
See publicationVisualizing NetFlow Data with Apache Kudu, Apache Impala (incubating), StreamSets Data Collector, and D3.js
Projects
-
Yardstick Spark
Technologies Used: Spark , Spark SQL CoreData Frame's, Spark Streaming, Scala, HDFS, Apache Ignite, Yardstick Tool, D3js, Tableau, AWS
Data Sets: 10 million twitter data records and 1 billion auto generated records.
Description:
Yardstick Apache Spark is a set of Apache Spark benchmarks written on top of Yardstick
framework.
I have successfully written Spark CoreRDD application to read auto generated 1
…Technologies Used: Spark , Spark SQL CoreData Frame's, Spark Streaming, Scala, HDFS, Apache Ignite, Yardstick Tool, D3js, Tableau, AWS
Data Sets: 10 million twitter data records and 1 billion auto generated records.
Description:
Yardstick Apache Spark is a set of Apache Spark benchmarks written on top of Yardstick
framework.
I have successfully written Spark CoreRDD application to read auto generated 1
billion records and compare with IgniteRDD in Yardstick framework to measure
performance of Apache Ignite RDD and Apache Spark RDD.
I have successfully written Spark DataFrame application to read from HDFS and analyze
10 million twitter records using Yardstick framework to measure performance of Apache
Ignite SQL and Apache Spark DataFrame.
I have successfully written Spark Streaming application to read streaming twitter data and
analyze twitter records in real time using Yardstick framework to measure performance of
Apache Ignite Streaming and Apache Spark Streaming.
Implemented test cases for Spark and Ignite functions using scala as language.
Hands-on experience in setting up 10 node Spark cluster on Amazon Web Service’s
using Spark EC2 script.
Implemented and D3.js Tableau charts to show performance difference between Apache
Ignite and Apache Spark.
Other creatorsSee project -
Comparative Analysis of Big Data Analytical Tools –(Hive,Hive on Tez,Impala,SparkQL,Apache Drill, BigQuery, PrestoDB running on the Google Cloud and AWS)
Technologies Used: Hive,Hive on Tez,Impala,SparkSQL, Apache Drill, BigQuery, PrestoDB, Hadoop, Cloudera
CDH, Hortonworks HDP, Google Cloud Platform, Amazon web service
Data set used: Twitter streaming data
a. Installation of Hive,Hive on Tez,Impala, SparkQL, Apache Drill, BigQuery, PrestoDB, Hadoop,
Cloudera CDH, Hortonworks HDP
b. Schema design for data sets on all Hive,Hive on Tez,Impala, SparkQL, Apache Drill, BigQuery,
PrestoDB
c. Query design for given data set
d…Technologies Used: Hive,Hive on Tez,Impala,SparkSQL, Apache Drill, BigQuery, PrestoDB, Hadoop, Cloudera
CDH, Hortonworks HDP, Google Cloud Platform, Amazon web service
Data set used: Twitter streaming data
a. Installation of Hive,Hive on Tez,Impala, SparkQL, Apache Drill, BigQuery, PrestoDB, Hadoop,
Cloudera CDH, Hortonworks HDP
b. Schema design for data sets on all Hive,Hive on Tez,Impala, SparkQL, Apache Drill, BigQuery,
PrestoDB
c. Query design for given data set
d. Debugging on Hive,Hive on Tez,Impala, SparkQL, Apache Drill, BigQuery, PrestoDB, Hadoop,
Cloudera CDH, Hortonworks HDP
e. Time Comparison of each Hive,Hive on Tez,Impala, SparkQL, Apache Drill, BigQuery, PrestoDB
f. Time comparison between different cloud platforms
g. Times metrics web based visualization design on google charts
Other creatorsSee project -
Truck Events Analysis
Technologies Used: Hadoop, HDFS, Hive, HBase, Kafka, Storm, RabbitMQ WebStormp, Google Maps
Data Sets Used: - New York City Truck Routes from NYC DOT. -Truck Events Data generated using a custom
simulator. - Weather Data, collected using APIs from Forcast.io. -Traffic Data, collected using APIs from
MapQuest.
a. Written a simulator to send/emit events based on NYC DOT data file.
b. Written Kafka Producer to accept/send events to Kafka Producer which is on Storm Spout
c…Technologies Used: Hadoop, HDFS, Hive, HBase, Kafka, Storm, RabbitMQ WebStormp, Google Maps
Data Sets Used: - New York City Truck Routes from NYC DOT. -Truck Events Data generated using a custom
simulator. - Weather Data, collected using APIs from Forcast.io. -Traffic Data, collected using APIs from
MapQuest.
a. Written a simulator to send/emit events based on NYC DOT data file.
b. Written Kafka Producer to accept/send events to Kafka Producer which is on Storm Spout
c. Written Storm topology to accept events from Kafka Producer and Process Events
d. Written Storm Bolt to Emit data into Hbase, HDFS, RabbitMQ WebStomp
e. Hive Queries to Map Truck Events Data , Weather Data, Traffic Data
Other creators -
Cimbal/MobApp Pay
Cimbal is a mobile promotion and payment network designed to increase business sales and deliver targeted deals to
consumers.
Technologies used: Hadoop MapReduce, Hbase , Spring Data Rest Web Service, CDH
Data Set : Users Payment Data
a. Written MapReduce programs to validate the data
b. Written more than 50 Spring Data Hbase rest API's in Java
c. Schema design on Hbase and cleaning data
d. Written Hive queries for analytic's on users data.
Other creatorsSee project -
Social Media Sentiment Analysis Update on each 4 hours (Positive Negative Review's)
Social Media Sentiment Analysis project gives positive negative reviews on our favorite words search, it is based on Twitter tweets, Facebook status's etc... it has been implemented using Social media API's with BigData Technologies like Flume to get social media stream data regularly into HDFS and Cassandra, MongoDB to store sentiment analyzed data. and in this project i am using apache oozie to develop work flow and it will be ruing on each 4 hours. this is project is done with my self…
Social Media Sentiment Analysis project gives positive negative reviews on our favorite words search, it is based on Twitter tweets, Facebook status's etc... it has been implemented using Social media API's with BigData Technologies like Flume to get social media stream data regularly into HDFS and Cassandra, MongoDB to store sentiment analyzed data. and in this project i am using apache oozie to develop work flow and it will be ruing on each 4 hours. this is project is done with my self interest.
-
Whole Genome Sequencing - Quality Check (QC)
See project-Entire program was developed in Hadoop MapReduce praradigm
-Performed custom Quality Check on genomic data to retain good quality reads
-Implemented novel features in the software:
-capability to handle sequencing-machine errors
-automatic detection of base-line PHRED score
-being platform agnostic and able to handle various file-formats i.e. from Illumina, 454 Roche, Complete Genomics, ABI Solid input format data -
Whole Genome Sequencing - Sequence Alignment
See project- Developed a Hadoop MapReduce program to perform sequence alignment on NGS data.
- The MapReduce program implements algorithms such as Borrows-Wheeler Transform (BWT), Ferragina-Manzini Index (FMI), Smith-Waterman dynamic programming algorithm using Hadoop distributed cache. -
Seismic Data Server & Repository (SDSR™)
Our Seismic Data Server & Repository solves the problem of delivering, on demand, precisely cropped SEG-Y files for instant loading at geophysical interpretation workstations anywhere in the network. Based on Hadoop file storage, Hbase™ and MapReduce technology, the Seismic Data Server brings fault-tolerant petabyte-scale store capability to the industry. Seismic Data Server supports post-stack traces now with pre-stack support to be released shortly.
Other creatorsSee project -
DDSR (Drilling Data Search and Repository):
This project aims to provide analytics for Oil and Gas exploration data. This DDSR repository builded by using HBase, Hadoop and its sub projects. We are collecting thousands of wells data from across the globe. This data is stored in Hbase and Hive by using Hadoop MapReduce jobs. On top of this data we are building analytics for search and advanced search.
Other creatorsSee project -
E Commerce (Obesessory.com)
-
See projectTech: Hsearch(Hbase+lucene), Hbase, Cassandra, Hive, Spark, Hadoop, MapReduce,Amazon Web
Service, Linode, CDN, Scala, Java
Data: Affilats feeds Rakuten, CJ, Affiliate window, WebGains.
Pre-Processing:
Crawling of 100+ sites Data using Nutch
Fashion based ontology maintenance
Using Spark, Scala:
to Enriched given data using Fashion Ontology
to Validation/Normalizing the data
Designed schema, Modeling the data and Written program to store all validated data in…Tech: Hsearch(Hbase+lucene), Hbase, Cassandra, Hive, Spark, Hadoop, MapReduce,Amazon Web
Service, Linode, CDN, Scala, Java
Data: Affilats feeds Rakuten, CJ, Affiliate window, WebGains.
Pre-Processing:
Crawling of 100+ sites Data using Nutch
Fashion based ontology maintenance
Using Spark, Scala:
to Enriched given data using Fashion Ontology
to Validation/Normalizing the data
Designed schema, Modeling the data and Written program to store all validated data in Cassandra
Spring Data Cassandra Programs for Validation/Normalizing/Enriching and REST API to Develop UI Based manual QA Validation.
Used SPARQL, Scala to running QA based SQL queries.
Indexing:
MR Programs on Hbase:
To Standardize the Input Merchants data
To Upload images to RackSpace CDN
To Index the given Data sets into Hsearch
To MR Programs on Hbase to extract the color information from Images including density
To MR Programs on Hbase to Persist the Data on Hbase tables
above MR jobs will run based on timing and bucketing
Color-Obsessed:
Using Image color and density data
User will be allowed to select 1,2.. colors with different densities and result will be a list of
products where each product image contains all give colors with exact density
Written Hbase Spring rest web service for Color Obsessed Search API
Post-Processing:
Setting up the Spark Streaming and Kafka Cluster
Developed a Spark Streaming_Kafka App to Process Hadoop Jobs Logs
Kafka Producer to send all slaves logs to Spark Streaming App
Spark Streaming App to Process the Logs with given rules and produce the Bad Images,Bad
records, Missed Records etc
Spark Streaming App collect user actions data from front end
Kafka Producer based Rest API to collect user events and send to Spark Streaming App
Hive Queries to Generate Stock Alerts, Price Alerts, Popular Products Alerts,
New Arrivals for each user based on given likes, favorite,shares counts information
Working on SparkMLlib for Recommendations, Coupons Recommendations, Rules
Engine
Languages
-
English
Professional working proficiency
-
Telugu
Native or bilingual proficiency
-
Kannada
Native or bilingual proficiency
-
Tamil
Limited working proficiency
Recommendations received
-
LinkedIn User
4 people have recommended Sandish Kumar
Join now to viewOther similar profiles
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content