Posts

Dealing with NP-Hard Problems: An Introduction to Approximation Algorithms

Image
This is just a quick overview on approximation algorithms. It is a broad topic to discuss. For more info rmation go to References. The famous NP-Complete class is known for its possible intractability. NP means non deterministic polynomial and for a problem to be NP-Complete it has to be   NP  (verified in polynomial time) and   NP-Hard (as hard as any other problem in the NP class). Among the several important problems that are NP-Complete or NP-Hard (on its optimization form) we can name the Knapsack, the Travel Salesmen, and the Set Cover problem. Even though no efficient optimal solution might exist for NP-Complete problems we still need to address this issue due to the amount of practical problems existent in the NP-Complete class. Considering that even for medium volumes of data exponential brute-force is impractical, the option is abdicating the optimum solution as minimum as possible and pursuing an efficient algorithm.   Approximation a...

Overview of Digital Cloning

Image
Introduction The growth of the image processing and editing software availability has made it easy to manipulate digital images. With the amount of digital content being generated nowadays, developing  techniques to verify the authenticity and integrity of digital content might be essential to provide truthful evidences in a forensics case. In this context, copy-move is a type of forgery in which a part of an image is copied and pasted somewhere else in the same image . This forgery might be particularly challenging to discover due to properties like illumination  and noise matching on the source and the tampered regions. An example of copy-move forgery can be seen in picture 1. First we can see the original image, followed by the tampered one, and then a picture with the indication of the cloned areas. Several techniques have been proposed to solve this problem. The Block-based methods [1] divide an image in blocks of pixels and compare them to find a forgery. Ke...

Understanding Apache Hive

Image
   Introduction   BigData and Hive Apache Hive is a software application created to facilitate data analyses on Apache Hadoop. It is a Java framework that helps extracting knowledge from data placed on a HDFS cluster by providing a SQL-like interface to it. The Apache Hadoop platform is a major project on distributed computing and it is commonly assumed to be the best approach when dealing with BigData challenges. It is now very well established that great volume of data is produced everyday. Whether it is by system logs or by users purchases, the amount of information generated is such that previous existing Databases and Datawarehouses solutions don’t seem to scale well enough. The MapReduce programming paradigm was uncovered in 2004 as a new approach on processing large datasets. In 2005 its OpenSource version, Hadoop, was created by Doug Cutting. Although Hadoop is not set for substituting relational databases, it is a good solution for big...

Is there such a thing as "best" Recommender System algorithm?

I received emails from users asking which recommender system algorithm they should use . Usually people start looking for articles on which approach has a better performance, and once they find something convincing they start to implement it. I believe that the best recommender system depends on the data and the problem you have to deal with. With that in mind, I decided to publish here some pros and cons for each recommender type (collaborative, content and hybrid), so people can decide for their own what algoritms better suit their needs. I've already presented these approaches here , so if you know nothing about recommender systems, you can read it there first. Collaborative Filtering Pros Recommends diverse items to users, being innovative; Good practical results (read Amazon's article ); It is widely used, and you can find several OpenSource  implementations of it ( Apache Mahout ); It can be used on ratings from users on items; It can deal with video and...

Recommender Systems Online Free Course on Coursera

I already talked about Coursera's great courses here . There is a new course on Recommender Systems starting in September: https://www.coursera.org/course/recsys I don't know how it is going to be, but based on the courses I've done so far, it looks good.

Apache Hive .orig test file and "#### A masked pattern was here ####"

Just a quick information about something in Hive. If you ever typed: $ ant clean package test to run Apache Hive unit tests, you may have seen that Hive sometimes creates two output files . If you run for example: $ ant test -Dtestcase=TestCliDriver -Dqfile=alter5.q Hive sometimes generates a alter5.q.out and a alter5.q.out.orig : build/ql/test/logs/clientpositive/alter5.q.out build/ql/test/logs/clientpositive/alter5.q.out.orig This happens because Hive uses a method to mask any local information, as local time, or local path, with the following sentence: #### A masked pattern was here #### So, if you check your .q.out file it should have a bunch of this sentence above covering several local information . This information needs to be covered so that the tests outputs are the same in all computers. The .q.out.orig file has the original test output, with all the local information non covered. Out of curiosity, the method to mask the local patterns (private void...

BigData Free Course Online

Coursera offers several great online courses from the best universities around the world. The courses involve video lectures being released weekly, work assignments for the student, and reading material indications.  I had enrolled on this course about BigData a couple of months ago, and I confess I didn't have time to start doing it since last week. Once I started the course I was pleased with the content presented. They talk about important Data Mining algorithms for dealing with great amount of data such as PageRank . MapReduce and Distributed File Systems are also two very well explained topics on this course. So, for those who want to know more about computing related to BigData this course is certainly recommended! https://www.coursera.org/course/bigdata PS: The course is being offered since march, and its inscriptions period must soon be over. But keep watching the course page, because they open new courses often.